Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

0. Setup: your workspace and engine

In this chapter

  • Create a personal workspace with the book's code, a Python environment and PyTorch.
  • Learn the loop you will repeat every chapter: run a test, implement, run it again.
  • Optionally prepare the GPU, C++ and Rust toolchains.

Time: about 1 hour. GPU: optional.

The workspace

Everything you do happens in one folder: your copy of the book’s code. It contains:

inference-zero-to-hero-code/
├── engine/      YOUR engine. Every function you'll write is there, with its signature
│                and docstring; the body says  raise NotImplementedError("TODO(Chapter 5) ...")
├── izh/         the complete reference engine: the answer key, same module and function names
├── tests/       one milestone test file per chapter (test_ch05_attention.py, ...)
├── run.py       a runnable demo for each chapter (python run.py --help)
├── data/        training text and small datasets (no downloads needed)
├── rust/  cpp/  the Rust and C++ tracks
└── models/      checkpoints you download later (created when you need it)

engine/ and izh/ have identical structure. That’s deliberate. When you’re stuck, you can open the same file in izh/ and read the reference. When you want to skip ahead, every test and demo can run against the reference instead of your engine:

IZH_IMPL=izh pytest tests/test_ch05_attention.py   # test the reference
python run.py attention --impl engine              # run a demo with YOUR engine

1. Install uv

uv manages Python versions and packages quickly and reproducibly. These commands are for Linux and macOS (Windows users: use WSL2):

curl -LsSf https://astral.sh/uv/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
uv --version

uv can download Python for you, so you don’t need a separate Python installation.

2. Create your working copy

Download the code archive (or, from a clone of the book’s repository, use docs/inference-zero-to-hero/src/code-examples.zip) and extract it into a fresh folder:

mkdir -p "$HOME/izh" && cd "$HOME/izh"
unzip -n ~/Downloads/code-examples.zip
cd inference-zero-to-hero-code
ls          # engine  izh  tests  run.py  data  rust  cpp ...

unzip -n never overwrites existing files, so re-extracting a newer archive won’t destroy your work. Put the folder under version control right away. Your engine is real code, and you’ll want its history:

git init && git add -A && git commit -m "Starting point"

3. Create the Python environment

uv venv --python 3.12
source .venv/bin/activate          # do this in every new terminal

Now install PyTorch. Choose one route.

# Let uv pick the CUDA build that matches your driver, then verify below.
uv pip install torch numpy pytest --torch-backend=auto
uv pip install torch numpy pytest --torch-backend=cpu
# macOS on M-series: PyTorch's default build includes the MPS backend.
uv pip install torch numpy pytest

Check that it worked, and that a GPU is really doing work if you have one:

python - <<'PY'
import torch
print("torch", torch.__version__, "| CUDA available:", torch.cuda.is_available())
if torch.cuda.is_available():
    print(torch.cuda.get_device_name(), torch.cuda.get_device_capability())
    x = torch.ones(4, device="cuda")
    print((x + x).cpu())        # tensor([2., 2., 2., 2.])
PY

Important

DGX Spark / GB10. Spark is an ARM64 machine with a Blackwell GPU of compute capability 12.1 (sm_121). Use a PyTorch wheel built for aarch64 with CUDA 13. --torch-backend=auto normally finds it. CPU and GPU share one pool of LPDDR5x memory (128 GB), which matters for the capacity planning in Chapters 18, 20 and 30.

4. The loop you will repeat every chapter

Run the first milestone’s tests. They fail, because you haven’t written anything yet:

pytest tests/test_ch03_autograd.py
E   NotImplementedError: TODO(Chapter 3): implement __add__ in engine/autograd.py

That message is your to-do list. Each chapter’s Build it section tells you which functions to implement and gives hints. You edit engine/<module>.py, rerun the tests, and repeat until they’re green:

pytest tests/test_ch03_autograd.py      # 5 passed
git commit -am "Chapter 3 milestone"

Then see your code do something:

python run.py autograd --impl engine

Tip

pytest -x stops at the first failure; pytest -k name runs only matching tests; pytest --pdb drops into a debugger where a test fails. The tests are short and readable, so open them: they are the specification.

Try it now with the reference, to confirm your environment works end to end:

IZH_IMPL=izh pytest -q        # ~126 passed, 12 skipped (the GPU-only tests) on a CPU
python run.py tensors

5. Packages you’ll add later

Only PyTorch is needed until Chapter 9. Later chapters ask you to install more, with the command shown at the point of use:

FromPackagesWhy
Chapter 9uv pip install -r optional-requirements.txttokenizers and chat templates (transformers), comparing your models with independent implementations
Chapter 11CUDA Toolkit, uv pip install ninja setuptoolscompiling your own CUDA kernels
Chapter 14triton (included with CUDA PyTorch; uv pip install triton on CPU)writing kernels in Python
Chapter 21uv pip install -r workflow-requirements.txtfine-tuning with PEFT

6. Models you’ll download later

The engine never downloads anything by itself. When a chapter needs a real checkpoint, it shows an explicit command like this one, which saves a pinned snapshot into models/:

uv pip install -r optional-requirements.txt
hf download Qwen/Qwen3-0.6B --local-dir models/Qwen3-0.6B
ChapterModelDiskPurpose
1, 17-23Qwen/Qwen3-0.6B1.5 GBthe dense model your engine runs
9openai-community/gpt20.5 GByour from-scratch GPT, loaded with real weights
26Qwen/Qwen3-1.7B (optional)4 GBa target for speculative decoding
27Qwen/Qwen3-30B-A3B (optional, quantized)17-60 GBa real mixture of experts
30Qwen/Qwen3.8-Flash-Next (optional)about 360 GBthe capstone at full scale

The optional large models are never required to pass a milestone. Every milestone is verified on small configurations, against independent reference implementations.

7. Optional: C++ and Rust

Every core mechanism in the book also exists in C++ and Rust, shown in tabs. The Rust track grows into a CPU engine that loads real Qwen3 weights. To follow along:

# C++17 compiler and CMake (Ubuntu: sudo apt install build-essential cmake)
cmake -S cpp -B build/cpp -DCMAKE_BUILD_TYPE=Release
cmake --build build/cpp -j
build/cpp/izh all          # self-checking demos, one per chapter
# Stable Rust from https://rustup.rs
cd rust
cargo test --release       # unit tests + parity with the Python reference
cargo run --release -- generate --model-dir ../models/Qwen3-0.6B --new-tokens 32

Appendix B describes both tracks and the optional NVIDIA Rust GPU toolchains.

When something goes wrong

SymptomFix
No module named torchsource .venv/bin/activate in this terminal; check which python
torch.cuda.is_available() is False on a GPU machineReinstall with --torch-backend=cu130 (or your CUDA version); check nvidia-smi
No module named engine when running a scriptRun commands from the workspace root, the folder that contains engine/
Triton tests are slow on a CPUExpected: the interpreter is about 100x slower than a GPU. Use pytest -k to run one
nvcc missing (Chapter 11)Install the CUDA Toolkit; a CUDA PyTorch wheel does not include the compiler
Out of memoryUse smaller batch or context; on Spark, check other processes with nvidia-smi

Record uv pip freeze and nvidia-smi output alongside any result you want to keep. Next: the big picture.