B. The C++ and Rust tracks
Python is the book’s main language because it’s where models, kernels (Triton) and tools meet. But an inference engine is systems software, and seeing each mechanism without PyTorch underneath is the best test that you understand it. Throughout the book, the tabs next to the Python code show the same mechanism in C++ and Rust. This appendix explains what each track contains and how to build it.
What each track covers
| chapter | mechanism | C++ (cpp/izh.hpp) | Rust (rust/src/) |
|---|---|---|---|
| 2 | strided tensors, views, matmul | strided, matmul | tensor.rs |
| 3 | scalar autograd | value | autograd.rs |
| 4 | byte-level BPE | bpe | bpe.rs |
| 5, 15 | attention; online softmax | attention, online | attention.rs |
| 6, 17 | LayerNorm, RMSNorm, GELU, SiLU, RoPE | norms, activations, rope | ops.rs |
| 8 | sampling (temperature, top-k, top-p) | sample | sampling.rs, rng.rs |
| 9 | safetensors, BF16 | safetensors | safetensors.rs |
| 11-13 | CUDA kernels | kernels.cu | rust-cuda/ (optional) |
| 16 | KV cache | cache | kv_cache.rs |
| 17-18 | a complete Qwen3 engine on the CPU | qwen3.rs, main.rs | |
| 20 | groupwise INT4 and packing | quant | quant.rs |
| 25 | block tables | paged | paged.rs |
| 26 | speculative acceptance | accept | speculative.rs |
| 28 | delta-rule step | delta | delta.rs |
| 40 | quantized CPU mat-vec and SIMD dispatch | cpp/izh_cpu.hpp: activation quantization, Q4/Q8 rows, runtime ISA selection | simd.rs: Q4 GEMV and x86 AVX2/scalar dispatch |
The C++ track has small, readable companions in izh.hpp and Chapter 40’s quantized kernels in izh_cpu.hpp, each exercised by a self-checking demo. The Rust track grows into a real engine: from Chapter 18 on, it loads Qwen3-0.6B (or any dense Qwen3) from the downloaded safetensors and generates text on the CPU.
C++
Requirements: a C++17 compiler and CMake 3.18 or newer.
cmake -S cpp -B build/cpp -DCMAKE_BUILD_TYPE=Release
cmake --build build/cpp -j
build/cpp/izh all # every chapter's demo
build/cpp/izh 15 # just Chapter 15's
Each demo checks its results and prints one line, for example:
ch15 online softmax equals dense attention for every tile size ok
ch20 int4 group quantization, packing, error bound ok
ch26 rejection correction reproduces p ok
cpp/kernels.cu is a standalone CUDA program, with no PyTorch, for Chapter 11’s first kernels (vector add and naive matmul), checked against a CPU computation. Build it with the CUDA Toolkit:
cmake -S cpp -B build/cpp -DBUILD_CUDA=ON && cmake --build build/cpp -j
build/cpp/izh_cuda
The full set of Chapters 11-13 (vector add, naive and tiled matmul, warp and block reductions, row sums, RMSNorm) lives in engine/kernels/cuda_ops.cu, compiled as a PyTorch extension by engine/kernels/cuda.py with torch.utils.cpp_extension.load, so the tests can check every kernel against PyTorch.
Rust
Requirements: stable Rust (from rustup.rs). The crate’s only dependency is serde_json, for safetensors headers and config.json.
cd rust
cargo test --release
The tests are of two kinds. Unit tests in each module check the mechanisms (the same worked examples as the book). tests/fixtures.rs loads two tiny Qwen3 checkpoints, one in F32 and one in BF16, written by the Python reference (python rust/make_fixtures.py regenerates them), and checks that the Rust engine reproduces the reference’s logits at every position to $10^{-4}$ and its greedy tokens exactly.
Running Qwen3 on the CPU
The engine works on token IDs. A small Python helper handles the text boundary with the checkpoint’s own tokenizer:
cd rust
IDS=$(python tokenize_ids.py encode --model-dir ../models/Qwen3-0.6B "Why is the sky blue?") # chat template; --raw for plain text
OUT=$(cargo run --release -- generate --model-dir ../models/Qwen3-0.6B --ids "$IDS" --new-tokens 64)
python tokenize_ids.py decode --model-dir ../models/Qwen3-0.6B "$OUT"
cargo run --release -- bench --model-dir ../models/Qwen3-0.6B # tokens/s
How it’s built, in qwen3.rs:
- Weights stay in BF16 in memory, exactly as on disk. The mat-vec widens each BF16 value to F32 inside the dot product, so memory traffic is 2 bytes per weight: the decode ceiling of Chapter 10 applies directly. Compare your measured tokens/s with your machine’s memory bandwidth divided by 1.19 GB.
- The mat-vec splits output rows across threads with
std::thread::scope. Decode is memory-bound, so a few threads saturate the memory bandwidth; more don’t help. - The KV cache is a preallocated
Vec<f32>per layer (kv_cache.rs), as in Chapter 16.
Chapter 40 adds simd.rs, with scalar and x86 AVX2 Q4 dots, checked against the scalar result and an error bound against floating-point weights. cpp/izh_cpu.hpp additionally supplies Q4/Q8 paths with AVX2, AVX-512 VNNI and compile-enabled ARM dot-product support; the Python CPU extension uses that layout. ISA coverage depends on the machine and build flags (Appendix F). Further work: wire quantized checkpoint storage into the native Qwen3 loader, add broader SIMD coverage and use the tokenizers crate to remove the Python helper.
Part VIII’s scheduler, HTTP server, model registry and distributed orchestration live in Python. These native tracks teach the underlying CPU mechanisms; they are not alternate implementations of all 44 milestones.
Optional: NVIDIA’s Rust GPU toolchains
Two experimental NVIDIA projects compile Rust for GPUs. They change quickly, so the book pins exact revisions and keeps the examples small:
- cuda-oxide (
rust-cuda/oxide/): Rust kernels in the SIMT model, one thread’s view like CUDA C++. Chapter 11 shows its vector kernel next to the CUDA C++ and Triton versions.install_example.pycopies the example into a fresh checkout of the pinnedNVIDIA/cuda-rustrevision; thencargo oxide run izh_vecaddbuilds and runs it. - cuTile Rust (
rust-cuda/cutile/): a tile-based model closer to Triton, where each program owns a tile of the output.cargo run --releasein that folder builds against the pinnedNVlabs/cutile-rsrevision.
Both need a recent NVIDIA driver and CUDA Toolkit, and neither is required for any milestone.