0. Setup: your workspace and engine
In this chapter
- Create a personal workspace with the book's code, a Python environment and PyTorch.
- Learn the loop you will repeat every chapter: run a test, implement, run it again.
- Optionally prepare the GPU, C++ and Rust toolchains.
Time: about 1 hour. GPU: optional.
The workspace
Everything you do happens in one folder: your copy of the book’s code. It contains:
inference-zero-to-hero-code/
├── engine/ YOUR engine. Every function you'll write is there, with its signature
│ and docstring; the body says raise NotImplementedError("TODO(Chapter 5) ...")
├── izh/ the complete reference engine: the answer key, same module and function names
├── tests/ one milestone test file per chapter (test_ch05_attention.py, ...)
├── run.py a runnable demo for each chapter (python run.py --help)
├── data/ training text and small datasets (no downloads needed)
├── rust/ cpp/ the Rust and C++ tracks
└── models/ checkpoints you download later (created when you need it)
engine/ and izh/ have identical structure. That’s deliberate. When you’re stuck, you can open the same file in izh/ and read the reference. When you want to skip ahead, every test and demo can run against the reference instead of your engine:
IZH_IMPL=izh pytest tests/test_ch05_attention.py # test the reference
python run.py attention --impl engine # run a demo with YOUR engine
1. Install uv
uv manages Python versions and packages quickly and reproducibly. These commands are for Linux and macOS (Windows users: use WSL2):
curl -LsSf https://astral.sh/uv/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
uv --version
uv can download Python for you, so you don’t need a separate Python installation.
2. Create your working copy
Download the code archive (or, from a clone of the book’s repository, use docs/inference-zero-to-hero/src/code-examples.zip) and extract it into a fresh folder:
mkdir -p "$HOME/izh" && cd "$HOME/izh"
unzip -n ~/Downloads/code-examples.zip
cd inference-zero-to-hero-code
ls # engine izh tests run.py data rust cpp ...
unzip -n never overwrites existing files, so re-extracting a newer archive won’t destroy your work. Put the folder under version control right away. Your engine is real code, and you’ll want its history:
git init && git add -A && git commit -m "Starting point"
3. Create the Python environment
uv venv --python 3.12
source .venv/bin/activate # do this in every new terminal
Now install PyTorch. Choose one route.
# Let uv pick the CUDA build that matches your driver, then verify below.
uv pip install torch numpy pytest --torch-backend=auto
uv pip install torch numpy pytest --torch-backend=cpu
# macOS on M-series: PyTorch's default build includes the MPS backend.
uv pip install torch numpy pytest
Check that it worked, and that a GPU is really doing work if you have one:
python - <<'PY'
import torch
print("torch", torch.__version__, "| CUDA available:", torch.cuda.is_available())
if torch.cuda.is_available():
print(torch.cuda.get_device_name(), torch.cuda.get_device_capability())
x = torch.ones(4, device="cuda")
print((x + x).cpu()) # tensor([2., 2., 2., 2.])
PY
Important
DGX Spark / GB10. Spark is an ARM64 machine with a Blackwell GPU of compute capability 12.1 (
sm_121). Use a PyTorch wheel built foraarch64with CUDA 13.--torch-backend=autonormally finds it. CPU and GPU share one pool of LPDDR5x memory (128 GB), which matters for the capacity planning in Chapters 18, 20 and 30.
4. The loop you will repeat every chapter
Run the first milestone’s tests. They fail, because you haven’t written anything yet:
pytest tests/test_ch03_autograd.py
E NotImplementedError: TODO(Chapter 3): implement __add__ in engine/autograd.py
That message is your to-do list. Each chapter’s Build it section tells you which functions to implement and gives hints. You edit engine/<module>.py, rerun the tests, and repeat until they’re green:
pytest tests/test_ch03_autograd.py # 5 passed
git commit -am "Chapter 3 milestone"
Then see your code do something:
python run.py autograd --impl engine
Tip
pytest -xstops at the first failure;pytest -k nameruns only matching tests;pytest --pdbdrops into a debugger where a test fails. The tests are short and readable, so open them: they are the specification.
Try it now with the reference, to confirm your environment works end to end:
IZH_IMPL=izh pytest -q # ~126 passed, 12 skipped (the GPU-only tests) on a CPU
python run.py tensors
5. Packages you’ll add later
Only PyTorch is needed until Chapter 9. Later chapters ask you to install more, with the command shown at the point of use:
| From | Packages | Why |
|---|---|---|
| Chapter 9 | uv pip install -r optional-requirements.txt | tokenizers and chat templates (transformers), comparing your models with independent implementations |
| Chapter 11 | CUDA Toolkit, uv pip install ninja setuptools | compiling your own CUDA kernels |
| Chapter 14 | triton (included with CUDA PyTorch; uv pip install triton on CPU) | writing kernels in Python |
| Chapter 21 | uv pip install -r workflow-requirements.txt | fine-tuning with PEFT |
6. Models you’ll download later
The engine never downloads anything by itself. When a chapter needs a real checkpoint, it shows an explicit command like this one, which saves a pinned snapshot into models/:
uv pip install -r optional-requirements.txt
hf download Qwen/Qwen3-0.6B --local-dir models/Qwen3-0.6B
| Chapter | Model | Disk | Purpose |
|---|---|---|---|
| 1, 17-23 | Qwen/Qwen3-0.6B | 1.5 GB | the dense model your engine runs |
| 9 | openai-community/gpt2 | 0.5 GB | your from-scratch GPT, loaded with real weights |
| 26 | Qwen/Qwen3-1.7B (optional) | 4 GB | a target for speculative decoding |
| 27 | Qwen/Qwen3-30B-A3B (optional, quantized) | 17-60 GB | a real mixture of experts |
| 30 | Qwen/Qwen3.8-Flash-Next (optional) | about 360 GB | the capstone at full scale |
The optional large models are never required to pass a milestone. Every milestone is verified on small configurations, against independent reference implementations.
7. Optional: C++ and Rust
Every core mechanism in the book also exists in C++ and Rust, shown in tabs. The Rust track grows into a CPU engine that loads real Qwen3 weights. To follow along:
# C++17 compiler and CMake (Ubuntu: sudo apt install build-essential cmake)
cmake -S cpp -B build/cpp -DCMAKE_BUILD_TYPE=Release
cmake --build build/cpp -j
build/cpp/izh all # self-checking demos, one per chapter
# Stable Rust from https://rustup.rs
cd rust
cargo test --release # unit tests + parity with the Python reference
cargo run --release -- generate --model-dir ../models/Qwen3-0.6B --new-tokens 32
Appendix B describes both tracks and the optional NVIDIA Rust GPU toolchains.
When something goes wrong
| Symptom | Fix |
|---|---|
No module named torch | source .venv/bin/activate in this terminal; check which python |
torch.cuda.is_available() is False on a GPU machine | Reinstall with --torch-backend=cu130 (or your CUDA version); check nvidia-smi |
No module named engine when running a script | Run commands from the workspace root, the folder that contains engine/ |
| Triton tests are slow on a CPU | Expected: the interpreter is about 100x slower than a GPU. Use pytest -k to run one |
nvcc missing (Chapter 11) | Install the CUDA Toolkit; a CUDA PyTorch wheel does not include the compiler |
| Out of memory | Use smaller batch or context; on Spark, check other processes with nvidia-smi |
Record uv pip freeze and nvidia-smi output alongside any result you want to keep. Next: the big picture.