Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

1. The big picture: what happens when you press Enter

In this chapter

  • Run a real language model on your machine in ten lines of Python.
  • Take it apart: tokens, embeddings, layers, logits, sampling, and the loop that ties them together.
  • Meet the two phases of inference, prefill and decode, and the one number that limits decode speed.
  • Get a map of the engine you'll build and where each part of the book fits.

You will build

No engine code yet. You'll measure your machine's generation speed and predict it from first principles.

Time: 2-3 hours. GPU: optional.

A model in ten lines

Let’s start with the destination and work backwards. Install the libraries that read real checkpoints, and download a small but genuinely capable model, Qwen3-0.6B (about 1.5 GB):

uv pip install -r optional-requirements.txt
hf download Qwen/Qwen3-0.6B --local-dir models/Qwen3-0.6B

Now ask it something. Create experiments/ch01.py in the code root (the directory containing run.py), save the following example there, and run python -m experiments.ch01 from that root. Use this same script for the measurements and stretch exercises below:

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

path = "models/Qwen3-0.6B"
tokenizer = AutoTokenizer.from_pretrained(path)
model = AutoModelForCausalLM.from_pretrained(path, dtype=torch.bfloat16).eval()

messages = [{"role": "user", "content": "Explain what a GPU is in one sentence."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
ids = tokenizer(text, return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=40, do_sample=False)
print(tokenizer.decode(out[0, ids.shape[1]:], skip_special_tokens=True))

You’ll get an answer along these lines (exact wording varies with library version and hardware):

A GPU (Graphics Processing Unit) is a specialized processor designed to rapidly perform
many calculations in parallel, originally for rendering graphics.

That’s the whole user-facing experience. A library you didn’t write did the work: model.generate. By Chapter 18, that library will be replaced by your engine, and you’ll know what every line of it does and why. The rest of this chapter opens the box.

Tip

No download possible right now? Everything below also works with the tiny random model in the reference package: python run.py cache and python run.py fast exercise the same machinery. The text will be gibberish, but the mechanics are identical.

Step 1: text becomes token IDs

A model never sees characters. A tokenizer chops text into pieces from a fixed vocabulary and replaces each piece with its integer ID:

print(ids[0].tolist())
print([tokenizer.decode([i]) for i in ids[0]])

The output looks like this (middle IDs elided):

[151644, 872, 198, ..., 151645, 198, 151644, 77091, 198]
['<|im_start|>', 'user', '\n', 'Ex', 'plain', ' what', ' a', ' GPU', ' is', ..., '<|im_end|>', '\n', '<|im_start|>', 'assistant', '\n']

Notice three things. Common words are single tokens (' GPU', with its leading space). Rarer words may be split into pieces. And the chat template added special tokens like <|im_start|> (ID 151644) and <|im_end|> (ID 151645) that mark who is speaking. Qwen3’s vocabulary has 151,936 entries. You’ll build a tokenizer like this, trained on a short novel, in Chapter 4.

Step 2: IDs become vectors

The model’s first layer is a lookup table, the embedding, with one row of 1,024 numbers per vocabulary entry. Token 22670 (' GPU') becomes row 22670. The whole prompt becomes a matrix with one row per token: shape [tokens, 1024]. These rows are learned during training so that tokens used in similar ways end up with similar vectors.

Step 3: the vectors flow through layers

Qwen3-0.6B has 28 transformer layers. Each one does two things to every token vector:

  1. Attention lets each token gather information from the tokens before it. That’s how ' sentence' comes to “know” it’s part of a request about GPUs. (Chapter 5)
  2. The MLP transforms each token’s vector on its own, applying what the model learned. (Chapter 6)

Each layer adds its result to the vector it received (the residual stream), so information accumulates as it flows upward. After 28 layers, each position holds a vector that summarizes everything relevant so far.

Step 4: the last vector becomes a prediction

The final vector, at the last position, is multiplied by a [1024 × 151936] matrix (for Qwen3-0.6B, the same matrix as the embedding table, “tied”). That produces 151,936 scores called logits, one per vocabulary entry: how plausible each token is as the next one.

with torch.no_grad():
    logits = model(ids).logits[0, -1]          # scores for the token after the prompt
probs = torch.softmax(logits.float(), dim=-1)
top = probs.topk(5)
for p, i in zip(top.values, top.indices):
    print(f"{tokenizer.decode([i])!r:12} {p:.3f}")

A typical result (your probabilities will differ a little):

'A'          0.91
'GPU'        0.05
'The'        0.02
...

Softmax turned scores into probabilities that sum to 1. Here the model is quite sure the answer starts with “A”.

Step 5: choose, append, repeat

Sampling picks one token from that distribution. Always taking the most likely one is called greedy decoding. Then comes the trick that makes generation work: append the chosen token to the input and run the model again. Each step produces exactly one new token. A 40-token answer takes 40 trips through all 28 layers. This loop is the autoregressive generation loop, and it’s the heart of inference:

ids_so_far = ids
for _ in range(40):
    with torch.no_grad():
        logits = model(ids_so_far).logits[0, -1]
    next_id = logits.argmax()                       # greedy
    ids_so_far = torch.cat([ids_so_far, next_id.view(1, 1)], dim=1)
    if next_id == tokenizer.convert_tokens_to_ids("<|im_end|>"):
        break

Play with the whole pipeline here:

Explore: one generation step

Step through what happens to a short prompt. Each press advances one stage; after "append" the loop starts again with one more token.

Two phases with very different costs

The loop above is wasteful. Every iteration re-processes the whole prompt, even though nothing about the earlier tokens has changed. Real engines remember each layer’s intermediate results for earlier tokens (the KV cache, Chapter 16), which splits generation into two phases:

  • Prefill. Process all prompt tokens at once, filling the cache. Hundreds or thousands of tokens flow through each layer together, as one big matrix multiplication. GPUs are superb at this, and prefill is limited by arithmetic speed (it’s compute-bound).
  • Decode. Produce one token at a time. Each step pushes a single token vector through every layer, reusing the cache. There’s little arithmetic per step, but every weight of the model must still be read from memory. Decode is memory-bound.

The user sees these as two numbers. Time to first token (TTFT) is mostly prefill. Inter-token latency, how fast the words stream, is decode.

The most important number in this book

Here is the back-of-envelope estimate that drives almost every design decision in modern inference engines.

To produce one token, decode multiplies one vector by every weight matrix in the model, so it must read every weight from memory once. Qwen3-0.6B has about 0.6 billion parameters. In BF16, each parameter takes 2 bytes, so each token requires reading about 1.2 GB.

How fast can memory deliver bytes? That’s the memory bandwidth of your device:

DeviceMemory bandwidthUpper bound for Qwen3-0.6B (1.2 GB/token)
Laptop CPU (DDR5, 2 channels)~80 GB/s~65 tokens/s
DGX Spark (GB10, LPDDR5x)273 GB/s~230 tokens/s
RTX 4090 (GDDR6X)1,008 GB/s~840 tokens/s
H100 SXM (HBM3)3,350 GB/s~2,800 tokens/s

$$ \text{tokens per second} ;\lesssim; \frac{\text{memory bandwidth (bytes/s)}}{\text{bytes read per token}} $$

That’s a ceiling: real engines reach 60-90% of it. Compare it with the arithmetic. One token needs about 2 floating-point operations per parameter, roughly 1.2 GFLOP. Even a laptop CPU does that in a few milliseconds, and a GPU in microseconds. During decode, the processor mostly waits for memory.

This one inequality explains a remarkable amount of the field:

  • Quantization (Chapter 20) stores weights in 4 bits instead of 16. That’s 4x fewer bytes per token, so decode gets close to 4x faster.
  • Batching (Chapter 24) decodes many users’ tokens in the same step. The weights are read once and used for every user, so throughput rises almost for free.
  • Speculative decoding (Chapter 26) checks several guessed tokens in one pass over the weights.
  • Mixture of experts (Chapter 27) reads only a few experts per token. Flash-Next stores 176B parameters but reads about 6B per token.
  • KV-cache size (Chapters 16, 25, 28) adds to the bytes per token, which is why long contexts are slow and why linear attention exists.

Explore: the decode speed limit

Pick a device and a model, change the weight precision and batch size, and see the bandwidth ceiling on tokens per second. Larger batches share each weight read across more tokens.

What an inference engine does

An inference engine is the program between “a folder of weights” and “a fast, correct stream of tokens for many users”. Its jobs, and where you’ll build each one:

JobWhat it meansWhere
Represent tensors and computematrix multiplications, attention, normsParts I-II
Load weightsread checkpoint files, map tensor names, check shapesChapters 9, 18
Run the model correctlyexactly the architecture the weights were trained forChapters 6, 17, 27-30
Manage stateKV cache, recurrent state, positionsChapters 16, 25, 28
Run it fastkernels, fusion, CUDA graphs, quantizationPart III, Chapters 19-20
Choose tokensgreedy, temperature, top-p, stop rulesChapter 8
Serve many requestsbatching, scheduling, memory sharing, speculationPart VI

Production engines like vLLM, SGLang and llama.cpp are hundreds of thousands of lines, mostly hardware-specific kernels and model adapters. The core ideas fit in a few thousand lines, and those are the lines you’ll write.

The destination: Qwen3.8-Flash-Next

The capstone model is a deliberate stretch. Qwen3.8-Flash-Next (released August 2026) is a preview of the architecture behind Qwen4. Its 48 layers mix:

  • Gated DeltaNet linear attention in 36 layers, which keeps a fixed-size memory instead of a growing KV cache (Chapter 28);
  • Qwen Sparse Attention in the other 12 layers, where a small indexer picks which 2,048 past tokens each query reads (Chapter 29);
  • a 512-expert mixture in every layer, 10 experts per token plus a shared one (Chapter 27);
  • four parallel residual streams with gated reads and writes, and a 51-billion-parameter hashed n-gram memory (Chapter 30).

Every one of those components is a variation on ideas you’ll learn first in their simplest form. By Chapter 30, you’ll write the whole model from scratch and verify it against the official implementation.

Build it

There’s no engine code this week. Instead, build your intuition with a measurement.

  1. Time the Hugging Face generation above for 100 new tokens (call torch.cuda.synchronize() before stopping the clock if you’re on a GPU), and compute tokens per second.
  2. Look up your device’s memory bandwidth and compute the ceiling from the formula.
  3. Compute the ratio. Below 50% is normal for this naive loop. Chapter 19 gets you much closer to the ceiling.
  4. Run the reference engine’s demos to see prefill and decode separately:
python run.py cache      # cached vs uncached generation time as output length grows
python run.py fast       # tokens/s of the plain loop vs a decoder without host round trips

Stretch exercises

  1. ★ Change do_sample=False to do_sample=True, temperature=1.2. Run it three times. What changed, and what didn’t? Where: copy the example in A model in ten lines into experiments/ch01.py (create it); change its model.generate call.
  2. ★★ Measure TTFT separately from decode speed by timing a 1-token generation and a 101-token one with a 500-token prompt. Which phase dominates for short answers? For long ones? Where: the same experiments/ch01.py (create it), around model.generate.
  3. ★★ Load the model in FP32 (dtype=torch.float32) and repeat the speed measurement. Predict the slowdown from the bandwidth formula before you run it. Where: the same script’s AutoModelForCausalLM.from_pretrained call.

Check your understanding

  1. Why does generating a 40-token answer require at least 40 sequential passes through the model?
  2. Prefill of a 1,000-token prompt and decode of one token both read every weight once. Why is prefill not 1,000x slower than one decode step?
  3. A 7B-parameter model in BF16 runs on a 1 TB/s GPU. What’s the most tokens per second a single user can get? What if the weights were 4-bit?

Going deeper

  • BALLM Chapter 1 (pp. 1-16): what LLMs are and how they’re built and trained.
  • PMPP §20.1 (pp. 478-482): the decoder-only transformer from a systems angle, including the generation loop.
  • GPU Mode L1 (profiling and integrating CUDA kernels in PyTorch) for a first look at where the time goes.
  • Horace He, Making Deep Learning Go Brrrr From First Principles (2022): the compute-bound / memory-bound / overhead-bound framing used throughout this book.