Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

A. The engine, tests and commands

Everything in this appendix is in the code download (the code/ folder of the book’s repository). Chapter 0 sets it up.

Layout

engine/      YOUR engine: every module, class and signature, with the chapter's functions left as TODO
izh/         the complete reference engine (the answer key): same modules, same names
tests/       one milestone test file per chapter, run against engine/ by default
run.py       one demo command per chapter; --impl engine runs yours
serve.py     OpenAI-compatible server launcher; see Chapter 36
izh/serve/   unified scheduler, block manager, runner, API and adapter slots
izh/bench.py, izh/evaluation.py
             endpoint load generator and correctness/quality harnesses (Chapter 44)
data/        The Verdict, BALLM's 1,100 instructions (with fixed splits), steering prompts
rust/        the Rust track: CPU engine for Qwen3, tests against the Python reference
cpp/         the C++ track: header-only companions and standalone CUDA kernels
rust-cuda/   optional NVIDIA Rust GPU examples
finetune.py, quantize_checkpoint.py, edit_checkpoint.py, model_workflows.py
             Hugging Face-based workflows for real checkpoints (Chapters 20-23)

engine/ is generated from izh/ by tools/make_engine.py: every function whose docstring says “(Your engine: Chapter N)” keeps its signature and docstring and has its body replaced with raise NotImplementedError("TODO(Chapter N): ..."). Triton kernels marked that way become pass, and the CUDA kernels in cuda_ops.cu lose their bodies. Everything not marked is provided, so you write the ideas and skip the boilerplate.

Warning

Regenerating overwrites engine/. The book’s repository does this to keep the skeleton in sync with the reference; you never need to. If you do, commit your work first.

Where to make exercise changes

All exercise paths are relative to the extracted code download’s root: the directory containing run.py, engine/ and tests/. In this repository that directory is docs/inference-zero-to-hero/src/code/.

  • Engine milestones: fill in the named TODOs in engine/. izh/ is the complete reference used for comparison. A code tab that includes code/izh/... shows that reference; the matching engine/... file is your implementation target.
  • Stretch exercises: each Where note names the edit target. These extensions can add functions, classes or arguments beyond the milestone’s existing TODOs. A file marked new or create it is a suggested file to create, not a supplied starter.
  • Experiments: create experiments/chNN.py for a chapter’s measurements, plots or standalone comparisons. Import your completed engine modules there. From the code root, run, for example, python -m experiments.ch02; this keeps the code root on Python’s import path. Use a notebook in the same code root if you prefer. Save handwritten predictions and observations alongside it.
  • Tests: run the chapter’s milestone suite after engine changes. Add extension checks in a new tests/test_chNN_stretch.py, importing the relevant engine modules. Milestone tests cover the required implementation; they do not automatically validate an optional extension.
  • Native tracks: edit the explicitly named cpp/ or rust/src/ file. For a new CUDA operation in engine/kernels/cuda_ops.cu, add its host launcher and PYBIND11_MODULE entry in that same file; engine/kernels/cuda.py loads the extension. Existing cpp/kernels.cu is the standalone companion.

The Check your understanding questions are written answers; they require no engine edits. Appendix C contains answers and stretch-exercise hints.

The chapter loop

pytest tests/test_ch16_kv_cache.py                 # red: NotImplementedError("TODO(Chapter 16): ...")
$EDITOR engine/kv_cache.py                          # implement the TODOs
pytest tests/test_ch16_kv_cache.py                 # green
python run.py cache --impl engine                   # watch your code run
IZH_IMPL=izh pytest tests/test_ch16_kv_cache.py    # the reference passes the same tests

Later chapters build on earlier ones: Chapter 17’s tests use your Chapter 5 attention and Chapter 16 cache. If a late test fails in an early function, fix the early one; its own tests may not have covered the case.

Environment variables:

variableeffect
IZH_IMPL=izhtests use the reference instead of engine/
IZH_DEVICE=cputests run on the CPU even if a GPU exists
TRITON_INTERPRET=1Triton kernels run in the interpreter (set automatically without a GPU)

Test markers: gpu (skipped without CUDA), reference (needs transformers; compares with the official implementation), slow. Run pytest -m "not slow" for a quick pass.

Milestones by chapter

ch.fileimplement
2tensors.pycontiguous_strides, element_offset, broadcast_shapes, matmul_loops, linear
3autograd.pyValue.__add__, __mul__, __pow__, exp, log, relu, tanh, backward
4tokenizer.py, data.pypair_counts, merge_pair, BPETokenizer.train, _encode_chunk; windows
5attention.pysplit_heads, merge_heads, causal_attention
6gpt.pyGPT.__init__, GPT.forward (and the attention and block classes)
7train.pylm_loss, evaluate, train
8sampling.pysample, generate_stream
9safetensors_io.py, loaders.pyread_header, load_file; load_gpt2
10measure.pymeasure_wall, matmul_intensity, attainable_flops, decode_ceiling
11-13kernels/cuda_ops.cuadd_kernel, naive_matmul_kernel, tiled_matmul_kernel, warp_sum, block_sum, row_sum_kernel, rmsnorm_kernel
12tiling.pytiled_matmul
13numerics.pystable_softmax
14kernels/triton_basics.py, triton_matmul.pyadd_kernel, softmax_kernel, add_rmsnorm_kernel, matmul_kernel
15attention.py, kernels/triton_flash.pyonline_attention; flash_fwd_kernel
16kv_cache.pyKVCache.update, KVCache.truncate
17qwen3.pyrope_cos_sin, apply_rope, RMSNorm.forward and the model’s forward methods
18loaders.py, engine.pyload_qwen3; LLM.stream
19fast.py, kv_cache.pysample_on_device, FastDecoder._step, generate; StaticKVCache.update
20quant.py, kernels/triton_quant.pyquantize_int8_rows, quantize_groupwise, QuantLinear.dequantized_weight, forward; w4a16_kernel
21lora.py, sft.pyinstruction_batch; last_token_logits
22lora.pyLoRALinear.forward, LoRALinear.merged
23steering.pydirection_from_means, project_out
24scheduler.pyScheduler.plan, ContinuousBatchingEngine.step
25paged.py, kernels/triton_paged.pyBlockAllocator.allocate, release; PagedKVCache.reserve, update, fork; paged_decode_kernel
26speculative.pyaccept_or_correct, speculative_generate
27moe.pyTopKRouter.forward, Experts.forward_grouped
28gdn.pyrecurrent_gated_delta_rule, chunk_gated_delta_rule, causal_conv1d, GatedDeltaNet.forward
29sparse.pyQSAIndexer.forward, GatedAttention.forward
30flashnext.pyGatedResidual.read, NGramEmbedding.shifted, forward, PLELayer.forward, FlashNextLayer.forward, FlashNext.step, load_flashnext
31serve/batch.py, blocks.py, scheduler.py, model.py, engine.pyragged batches, hash chains, paging, token-budget scheduling and the unified step loop
32kernels/triton_unified.py, serve/triton_backend.pyunified/split-KV attention, merging, quantized KV reads and writes
33serve/model.py, graphs.py, engine.py, kernels/triton_fused.pyprojection/MoE fusion, graph buckets and overlapping scheduling
34serve/sampler.py, beam.py, structured.pybatched sampling, beam search, regex/schema constraints
35serve/tokenizer.py, chat.pytokenizer loading, incremental UTF-8, stops and chat/tool boundaries
36serve/async_engine.py, api.pycore loop, async dispatch, generation and request preparation
37serve/spec.pyspeculation within the paged engine loop
38formats/gguf.py, ggml_quants.py, runtime.py, kernels/triton_formats.pyGGUF I/O, GGML decoding/repacking and affine matmul
39formats/gptq.py, awq.py, fp8.py, mx.py, trellis.pyquantized checkpoint layouts, quantization references and trellis decoding
40platform.py, offload.py, kernels/cpu.py, metal.pyplatform policies, CPU kernels and weight/layer offload
41parallel.pytensor/pipeline/expert parallelism, routing, KV transfer and ring attention
42models.pyconfigurable decoders, scaled RoPE, sliding windows, MLA and grouped routing
43multimodal.py, pooling.py, serve/adapters.py, runner.py, kernels/triton_lora.pyimage rows/M-RoPE, pooling, adapter slots and gathered shrink/expand kernels
44bench.py, evaluation.pyarrivals, stream clocks, goodput, teacher-forced quality and a task runner

rg "TODO\(Chapter" engine/ lists what’s left. Chapters 0, 1 and 45 have no implementation milestone.

run.py commands

Every command runs on a CPU unless noted; --impl engine runs your code; python run.py <command> --help lists the options.

commandchapterwhat it shows
tensors2shapes, strides, views and a gradient
autograd3a tiny MLP trained with your scalar autograd
bpe4a byte-level BPE trained on The Verdict
attention5a causal attention weight matrix
gpt6a GPT and its parameter count (and GPT-2 small’s)
train7training on The Verdict, with train and validation loss
generate8the trained model under several decoding policies
gpt29real GPT-2 weights through your GPT (--model-dir)
profile10wrong and right GPU timing, and a profiler trace
kernels11-15custom CUDA and Triton kernels against PyTorch
cache16cached against uncached generation time
chat18, 30a real checkpoint through your engine (--model-dir; Flash-Next: --expert-bits 4 --ngram-mmap)
fast19the plain loop against the static-buffer decoder (graph mode on CUDA)
quant20quantizing a model’s linear layers: size and logit error
classify21a GPT turned into a classifier (--data for SMS spam)
sft21instruction tuning (GPT-2 with --model-dir, or a GPT from scratch)
lora22LoRA on the sft model, then merging
steer23the logit lens and a steering direction
batch24continuous batching against one request at a time
paged25paged memory, prefix sharing and copy-on-write forks
speculate26speculative decoding with perfect and early-exit drafts
moe27router load, dispatch strategies, parameters of real MoEs
linear28associative recall, chunked against recurrent
sparse29top-k error and the indexer’s savings
flashnext30the real model’s memory plan and a tiny instance
core31mixed ragged scheduling, prefix reuse and memory pressure
backends32attention backends and split-KV comparisons
overhead33projection fusion, graph buffers and async scheduling
guided34batched sampling, constrained JSON and beam search
text35tokenizer parity, UTF-8 streaming, stops and chat parsing
serve36local random-model HTTP/SSE demo and client session
drafters37n-gram, draft-model and Medusa serving
gguf38GGUF export/load and quantized serving
quantformats39checkpoint layouts and quality comparisons
cpu, offload40quantized CPU mat-vec and expert-cache traffic
dist41distributed greedy parity and communication traffic
models42real configurations’ KV-memory arithmetic
features43a tiny vision prompt, mixed adapter requests and pooling
benchmark44worked SLO arithmetic and independent greedy parity; no performance claim

Server and measurement commands

Install serve-requirements.txt for HTTP, tokenizer and benchmark labs; Pillow is optional for image data URIs.

python serve.py --model-dir models/Qwen3-0.6B --device cpu --dtype fp32
python serve.py --help
python -m izh.bench --help
python -m izh.evaluation --help

Chapter 36 documents generation/API settings, Chapter 43 the preloaded --lora NAME=DIR adapters and embedding/reranking hooks, and Chapter 44 the complete workload/launch recipes and service limits. engine has the same module CLIs after its TODOs are implemented.

Workflows for real checkpoints

scriptchapterdoes
quantize_checkpoint.py20bitsandbytes NF4 or LLM.int8() against BF16: response loss, KL, bytes, reload check
finetune.py21-22full, LoRA or QLoRA fine-tuning of Qwen3 with response-only loss, manifest, reload
edit_checkpoint.py23fit and ablate a residual direction; loss and KL against the original
model_workflows.py20, 22, 23small self-contained labs: quantization trade-offs, LoRA, activation editing

They use Hugging Face Transformers for the model so that the chapter can focus on data, measurement and evaluation; your own engine implements the same ideas from scratch elsewhere in the chapter.