A. The engine, tests and commands
Everything in this appendix is in the code download (the code/ folder of the book’s repository). Chapter 0 sets it up.
Layout
engine/ YOUR engine: every module, class and signature, with the chapter's functions left as TODO
izh/ the complete reference engine (the answer key): same modules, same names
tests/ one milestone test file per chapter, run against engine/ by default
run.py one demo command per chapter; --impl engine runs yours
serve.py OpenAI-compatible server launcher; see Chapter 36
izh/serve/ unified scheduler, block manager, runner, API and adapter slots
izh/bench.py, izh/evaluation.py
endpoint load generator and correctness/quality harnesses (Chapter 44)
data/ The Verdict, BALLM's 1,100 instructions (with fixed splits), steering prompts
rust/ the Rust track: CPU engine for Qwen3, tests against the Python reference
cpp/ the C++ track: header-only companions and standalone CUDA kernels
rust-cuda/ optional NVIDIA Rust GPU examples
finetune.py, quantize_checkpoint.py, edit_checkpoint.py, model_workflows.py
Hugging Face-based workflows for real checkpoints (Chapters 20-23)
engine/ is generated from izh/ by tools/make_engine.py: every function whose docstring says “(Your engine: Chapter N)” keeps its signature and docstring and has its body replaced with raise NotImplementedError("TODO(Chapter N): ..."). Triton kernels marked that way become pass, and the CUDA kernels in cuda_ops.cu lose their bodies. Everything not marked is provided, so you write the ideas and skip the boilerplate.
Warning
Regenerating overwrites
engine/. The book’s repository does this to keep the skeleton in sync with the reference; you never need to. If you do, commit your work first.
Where to make exercise changes
All exercise paths are relative to the extracted code download’s root: the directory containing
run.py, engine/ and tests/. In this repository that directory is
docs/inference-zero-to-hero/src/code/.
- Engine milestones: fill in the named TODOs in
engine/.izh/is the complete reference used for comparison. A code tab that includescode/izh/...shows that reference; the matchingengine/...file is your implementation target. - Stretch exercises: each Where note names the edit target. These extensions can add functions, classes or arguments beyond the milestone’s existing TODOs. A file marked new or create it is a suggested file to create, not a supplied starter.
- Experiments: create
experiments/chNN.pyfor a chapter’s measurements, plots or standalone comparisons. Import your completedenginemodules there. From the code root, run, for example,python -m experiments.ch02; this keeps the code root on Python’s import path. Use a notebook in the same code root if you prefer. Save handwritten predictions and observations alongside it. - Tests: run the chapter’s milestone suite after engine changes. Add extension checks in a
new
tests/test_chNN_stretch.py, importing the relevantenginemodules. Milestone tests cover the required implementation; they do not automatically validate an optional extension. - Native tracks: edit the explicitly named
cpp/orrust/src/file. For a new CUDA operation inengine/kernels/cuda_ops.cu, add its host launcher andPYBIND11_MODULEentry in that same file;engine/kernels/cuda.pyloads the extension. Existingcpp/kernels.cuis the standalone companion.
The Check your understanding questions are written answers; they require no engine edits. Appendix C contains answers and stretch-exercise hints.
The chapter loop
pytest tests/test_ch16_kv_cache.py # red: NotImplementedError("TODO(Chapter 16): ...")
$EDITOR engine/kv_cache.py # implement the TODOs
pytest tests/test_ch16_kv_cache.py # green
python run.py cache --impl engine # watch your code run
IZH_IMPL=izh pytest tests/test_ch16_kv_cache.py # the reference passes the same tests
Later chapters build on earlier ones: Chapter 17’s tests use your Chapter 5 attention and Chapter 16 cache. If a late test fails in an early function, fix the early one; its own tests may not have covered the case.
Environment variables:
| variable | effect |
|---|---|
IZH_IMPL=izh | tests use the reference instead of engine/ |
IZH_DEVICE=cpu | tests run on the CPU even if a GPU exists |
TRITON_INTERPRET=1 | Triton kernels run in the interpreter (set automatically without a GPU) |
Test markers: gpu (skipped without CUDA), reference (needs transformers; compares with the official implementation), slow. Run pytest -m "not slow" for a quick pass.
Milestones by chapter
| ch. | file | implement |
|---|---|---|
| 2 | tensors.py | contiguous_strides, element_offset, broadcast_shapes, matmul_loops, linear |
| 3 | autograd.py | Value.__add__, __mul__, __pow__, exp, log, relu, tanh, backward |
| 4 | tokenizer.py, data.py | pair_counts, merge_pair, BPETokenizer.train, _encode_chunk; windows |
| 5 | attention.py | split_heads, merge_heads, causal_attention |
| 6 | gpt.py | GPT.__init__, GPT.forward (and the attention and block classes) |
| 7 | train.py | lm_loss, evaluate, train |
| 8 | sampling.py | sample, generate_stream |
| 9 | safetensors_io.py, loaders.py | read_header, load_file; load_gpt2 |
| 10 | measure.py | measure_wall, matmul_intensity, attainable_flops, decode_ceiling |
| 11-13 | kernels/cuda_ops.cu | add_kernel, naive_matmul_kernel, tiled_matmul_kernel, warp_sum, block_sum, row_sum_kernel, rmsnorm_kernel |
| 12 | tiling.py | tiled_matmul |
| 13 | numerics.py | stable_softmax |
| 14 | kernels/triton_basics.py, triton_matmul.py | add_kernel, softmax_kernel, add_rmsnorm_kernel, matmul_kernel |
| 15 | attention.py, kernels/triton_flash.py | online_attention; flash_fwd_kernel |
| 16 | kv_cache.py | KVCache.update, KVCache.truncate |
| 17 | qwen3.py | rope_cos_sin, apply_rope, RMSNorm.forward and the model’s forward methods |
| 18 | loaders.py, engine.py | load_qwen3; LLM.stream |
| 19 | fast.py, kv_cache.py | sample_on_device, FastDecoder._step, generate; StaticKVCache.update |
| 20 | quant.py, kernels/triton_quant.py | quantize_int8_rows, quantize_groupwise, QuantLinear.dequantized_weight, forward; w4a16_kernel |
| 21 | lora.py, sft.py | instruction_batch; last_token_logits |
| 22 | lora.py | LoRALinear.forward, LoRALinear.merged |
| 23 | steering.py | direction_from_means, project_out |
| 24 | scheduler.py | Scheduler.plan, ContinuousBatchingEngine.step |
| 25 | paged.py, kernels/triton_paged.py | BlockAllocator.allocate, release; PagedKVCache.reserve, update, fork; paged_decode_kernel |
| 26 | speculative.py | accept_or_correct, speculative_generate |
| 27 | moe.py | TopKRouter.forward, Experts.forward_grouped |
| 28 | gdn.py | recurrent_gated_delta_rule, chunk_gated_delta_rule, causal_conv1d, GatedDeltaNet.forward |
| 29 | sparse.py | QSAIndexer.forward, GatedAttention.forward |
| 30 | flashnext.py | GatedResidual.read, NGramEmbedding.shifted, forward, PLELayer.forward, FlashNextLayer.forward, FlashNext.step, load_flashnext |
| 31 | serve/batch.py, blocks.py, scheduler.py, model.py, engine.py | ragged batches, hash chains, paging, token-budget scheduling and the unified step loop |
| 32 | kernels/triton_unified.py, serve/triton_backend.py | unified/split-KV attention, merging, quantized KV reads and writes |
| 33 | serve/model.py, graphs.py, engine.py, kernels/triton_fused.py | projection/MoE fusion, graph buckets and overlapping scheduling |
| 34 | serve/sampler.py, beam.py, structured.py | batched sampling, beam search, regex/schema constraints |
| 35 | serve/tokenizer.py, chat.py | tokenizer loading, incremental UTF-8, stops and chat/tool boundaries |
| 36 | serve/async_engine.py, api.py | core loop, async dispatch, generation and request preparation |
| 37 | serve/spec.py | speculation within the paged engine loop |
| 38 | formats/gguf.py, ggml_quants.py, runtime.py, kernels/triton_formats.py | GGUF I/O, GGML decoding/repacking and affine matmul |
| 39 | formats/gptq.py, awq.py, fp8.py, mx.py, trellis.py | quantized checkpoint layouts, quantization references and trellis decoding |
| 40 | platform.py, offload.py, kernels/cpu.py, metal.py | platform policies, CPU kernels and weight/layer offload |
| 41 | parallel.py | tensor/pipeline/expert parallelism, routing, KV transfer and ring attention |
| 42 | models.py | configurable decoders, scaled RoPE, sliding windows, MLA and grouped routing |
| 43 | multimodal.py, pooling.py, serve/adapters.py, runner.py, kernels/triton_lora.py | image rows/M-RoPE, pooling, adapter slots and gathered shrink/expand kernels |
| 44 | bench.py, evaluation.py | arrivals, stream clocks, goodput, teacher-forced quality and a task runner |
rg "TODO\(Chapter" engine/ lists what’s left. Chapters 0, 1 and 45 have no implementation milestone.
run.py commands
Every command runs on a CPU unless noted; --impl engine runs your code; python run.py <command> --help lists the options.
| command | chapter | what it shows |
|---|---|---|
tensors | 2 | shapes, strides, views and a gradient |
autograd | 3 | a tiny MLP trained with your scalar autograd |
bpe | 4 | a byte-level BPE trained on The Verdict |
attention | 5 | a causal attention weight matrix |
gpt | 6 | a GPT and its parameter count (and GPT-2 small’s) |
train | 7 | training on The Verdict, with train and validation loss |
generate | 8 | the trained model under several decoding policies |
gpt2 | 9 | real GPT-2 weights through your GPT (--model-dir) |
profile | 10 | wrong and right GPU timing, and a profiler trace |
kernels | 11-15 | custom CUDA and Triton kernels against PyTorch |
cache | 16 | cached against uncached generation time |
chat | 18, 30 | a real checkpoint through your engine (--model-dir; Flash-Next: --expert-bits 4 --ngram-mmap) |
fast | 19 | the plain loop against the static-buffer decoder (graph mode on CUDA) |
quant | 20 | quantizing a model’s linear layers: size and logit error |
classify | 21 | a GPT turned into a classifier (--data for SMS spam) |
sft | 21 | instruction tuning (GPT-2 with --model-dir, or a GPT from scratch) |
lora | 22 | LoRA on the sft model, then merging |
steer | 23 | the logit lens and a steering direction |
batch | 24 | continuous batching against one request at a time |
paged | 25 | paged memory, prefix sharing and copy-on-write forks |
speculate | 26 | speculative decoding with perfect and early-exit drafts |
moe | 27 | router load, dispatch strategies, parameters of real MoEs |
linear | 28 | associative recall, chunked against recurrent |
sparse | 29 | top-k error and the indexer’s savings |
flashnext | 30 | the real model’s memory plan and a tiny instance |
core | 31 | mixed ragged scheduling, prefix reuse and memory pressure |
backends | 32 | attention backends and split-KV comparisons |
overhead | 33 | projection fusion, graph buffers and async scheduling |
guided | 34 | batched sampling, constrained JSON and beam search |
text | 35 | tokenizer parity, UTF-8 streaming, stops and chat parsing |
serve | 36 | local random-model HTTP/SSE demo and client session |
drafters | 37 | n-gram, draft-model and Medusa serving |
gguf | 38 | GGUF export/load and quantized serving |
quantformats | 39 | checkpoint layouts and quality comparisons |
cpu, offload | 40 | quantized CPU mat-vec and expert-cache traffic |
dist | 41 | distributed greedy parity and communication traffic |
models | 42 | real configurations’ KV-memory arithmetic |
features | 43 | a tiny vision prompt, mixed adapter requests and pooling |
benchmark | 44 | worked SLO arithmetic and independent greedy parity; no performance claim |
Server and measurement commands
Install serve-requirements.txt for HTTP, tokenizer and benchmark labs; Pillow is optional for image data URIs.
python serve.py --model-dir models/Qwen3-0.6B --device cpu --dtype fp32
python serve.py --help
python -m izh.bench --help
python -m izh.evaluation --help
Chapter 36 documents generation/API settings, Chapter 43 the preloaded --lora NAME=DIR adapters and embedding/reranking hooks, and Chapter 44 the complete workload/launch recipes and service limits. engine has the same module CLIs after its TODOs are implemented.
Workflows for real checkpoints
| script | chapter | does |
|---|---|---|
quantize_checkpoint.py | 20 | bitsandbytes NF4 or LLM.int8() against BF16: response loss, KL, bytes, reload check |
finetune.py | 21-22 | full, LoRA or QLoRA fine-tuning of Qwen3 with response-only loss, manifest, reload |
edit_checkpoint.py | 23 | fit and ablate a residual direction; loss and KL against the original |
model_workflows.py | 20, 22, 23 | small self-contained labs: quantization trade-offs, LoRA, activation editing |
They use Hugging Face Transformers for the model so that the chapter can focus on data, measurement and evaluation; your own engine implements the same ideas from scratch elsewhere in the chapter.