Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Inference Zero To Hero

Build your own LLM inference engine from first principles, and understand every byte it moves.

When you send a message to a chatbot, a program reads billions of numbers from memory for every word it writes back. Somebody wrote that program. By the end of this book, you will have written one too. It will be a real inference engine that loads open-weight models from disk, runs them on your GPU (or CPU), streams answers, serves many users at once, and implements the architecture of Qwen3.8-Flash-Next: a 176-billion-parameter hybrid model with linear attention, sparse attention, a 512-expert mixture, hashed n-gram memory and multi-stream residuals.

You will also be able to change models: fine-tune them, attach LoRA adapters, quantize them, look inside them and edit what they compute.

Who this book is for

You are a working software engineer. You can read and write Python, you know what a function, a class and a loop are, and you are comfortable in a terminal. You do not need:

  • prior machine-learning experience: we start from “what is a gradient?”;
  • GPU or CUDA experience: we start from “what is a thread?”;
  • heavy math: high-school algebra plus a willingness to follow small worked examples. When we need a derivative or a matrix identity, we derive it on the page with numbers.

If you already know some of this, skim the early chapters, but do the engine milestones. Later chapters assume your engine works.

What you will build

The book is organized around one project: your engine. Every chapter adds a piece and ends with a milestone test suite you make pass.

PartYou buildVisible result
I. Neural networkstensors, an autograd engine, a BPE tokenizera network that learns; a tokenizer trained on a novel
II. A GPT from scratchattention, transformer blocks, training, sampling, a checkpoint loaderyour GPT writes text; real GPT-2 runs in your code
III. GPU programmingCUDA and Triton kernels, FlashAttentionkernels you wrote, measured against PyTorch’s
IV. A dense engineKV cache, Qwen3, engine v1, fast decode, quantizationchat with Qwen3 through your engine at near-hardware speed
V. Changing modelsfine-tuning, LoRA/QLoRA, steering and abliterationyour own adapters and model edits
VI. Servingcontinuous batching, paged KV cache, speculative decodingmany requests at once, sharing memory safely
VII. Frontier architecturesMoE, Gated DeltaNet, sparse attention, Flash-Nexta from-scratch Flash-Next that matches the official implementation
VIII. Putting the engine togetherunified batching/paging, kernels, HTTP, speculation, formats, offload, parallelism, more models, images and adaptersone serving loop, an API, and correctness/latency evidence with explicit limits

The book has 46 chapters (including setup), in eight parts. The reference implementation passes the milestone suites your engine will use; Appendix F records current counts and platform coverage. Its tiny FP32 Flash-Next agrees with the official Hugging Face implementation to about 1e-7. Real-checkpoint and production performance claims require separate validation.

How each chapter works

Every chapter follows the same rhythm, so you always know where you are:

  1. Why it matters: the problem, tied to the engine.
  2. Concepts: intuition first (a picture, an analogy, a small example worked with real numbers), then the precise version.
  3. See it: diagrams and interactive playgrounds you can poke.
  4. Code it: the real implementation, in tabs. Python is the main language. C++ and Rust versions sit one click away, and your choice is remembered:
def softmax(x):
    e = [math.exp(v - max(x)) for v in x]
    return [v / sum(e) for v in e]
void softmax(float* x, size_t n) {
    float m = *std::max_element(x, x + n), sum = 0;
    for (size_t i = 0; i < n; ++i) sum += (x[i] = std::exp(x[i] - m));
    for (size_t i = 0; i < n; ++i) x[i] /= sum;
}
#![allow(unused)]
fn main() {
fn softmax(x: &mut [f32]) {
    let max = x.iter().cloned().fold(f32::NEG_INFINITY, f32::max);
    let sum: f32 = x.iter_mut().map(|v| { *v = (*v - max).exp(); *v }).sum();
    x.iter_mut().for_each(|v| *v /= sum);
}
}
  1. Build it: an engine milestone. You implement the chapter’s functions in your engine/ package and run its tests. The book supplies all the scaffolding (file layout, I/O, plotting, test harness), so you only write the ideas.
  2. Stretch exercises: graded ★ to ★★★. Each Where note identifies the file to edit or a new experiment to create; Appendix A explains paths, imports and extension tests.
  3. Check your understanding: short questions, with answers in Appendix C.
  4. Going deeper: exact pointers into the source books, lectures and papers.

Callouts mark things worth stopping for:

Note

Background or a useful aside.

Tip

A practical shortcut or a debugging habit.

Warning

A pitfall that produces plausible-looking wrong answers.

Important

A platform-specific note (for example, DGX Spark), or something that changes how you should read the rest.

Pacing

At 10 to 12 hours a week, plan on roughly one chapter a week. The GPU chapters and the capstone take two. The full eight-part path is about a year, with something working at the end of every week. If you have more time, the milestones are designed to be done in one sitting each.

You can read the whole book without a GPU. The reference milestones have CPU paths, including Triton’s interpreter; CUDA kernels, graph capture and hardware-specific adapters need their target device. A GPU makes the performance chapters far more satisfying, though. The book was developed for an NVIDIA DGX Spark (GB10), and notes for that machine are marked.

Where this book comes from

The book draws on three excellent resources, and borrows from them freely:

  • Sebastian Raschka, Build a Large Language Model (From Scratch) (Manning, 2024), cited as BALLM. Parts I and II follow its path from tokens to a trained GPT and reuse its corpus and several of its examples.
  • Wen-mei Hwu, David Kirk and Izzat El Hajj, Programming Massively Parallel Processors, 5th edition, cited as PMPP. Part III follows its approach to CUDA, tiling, reductions and attention.
  • The GPU Mode lecture series (github.com/gpu-mode/lectures), cited as Ln, for profiling, Triton, FlashAttention, quantization and serving.

You never need them open beside this book, but each chapter ends with exact pointers for a second explanation. Appendix E maps every chapter to its sources.

Start with Setup.