6. The transformer block and a complete GPT
In this chapter
- The residual stream: why each layer adds to its input instead of replacing it.
- LayerNorm and the GELU MLP, the other two ingredients of a transformer block.
- Assembling embeddings, blocks and an output head into a GPT, and counting its parameters exactly.
You will build
engine/gpt.py: a GPT-2-compatible model. In Chapter 7 you train it from scratch; in Chapter 9 it runs OpenAI's real GPT-2 weights unchanged.
Time: 4-6 hours. GPU: not needed.
The shape of a GPT
A GPT (“generative pre-trained transformer”) is a stack of identical blocks between an input embedding and an output head:
token IDs [B, T]
│ token embedding + position embedding
▼
x [B, T, D] ──► block 1 ──► block 2 ──► ... ──► block L ──► final LayerNorm ──► head ──► logits [B, T, V]
Every block has the same shape in and out, [B, T, D], so blocks stack like Lego. GPT-2 small has $L = 12$ blocks of width $D = 768$; Qwen3-0.6B has 28 blocks of width 1,024. Once you’ve built one block, you’ve built the model.
The residual stream
Each block contains two branches, attention and an MLP. Each branch reads the current vector, computes something, and adds its result back:
$$ \begin{aligned} a &= x + \operatorname{Attention}(\operatorname{LN}_1(x)) \ y &= a + \operatorname{MLP}(\operatorname{LN}_2(a)) \end{aligned} $$
Picture the vector $x$ as a residual stream that flows from the embedding to the output. Each branch reads from it and writes a small update into it. This has three big consequences:
- Training works at depth. The gradient of $x + f(x)$ with respect to $x$ is $1 + f’(x)$. Even if a branch’s gradient is tiny, the “1” carries the signal straight back to early layers. Before residual connections (He et al., 2015), very deep networks barely trained.
- A block can do nothing. If both branches output zero, the block is the identity. A new block starts as a small perturbation of a working network, not a random scrambling of it. Your milestone test checks exactly this: zero the output projections and the block must return its input.
- Interpretability. Every branch reads and writes the same space, so you can inspect, edit or steer what’s in the stream. Chapter 23 does exactly that.
This “pre-norm” arrangement (normalize the input of each branch) is used by GPT-2 and every modern model. The original 2017 transformer normalized after the addition (post-norm), which is harder to train at depth.
LayerNorm: keep the numbers in range
As updates accumulate in the stream, its vectors can drift to very different scales from token to token. Layer normalization rescales each token vector, independently of every other token and every other example, to mean 0 and variance 1, then applies a learned per-feature scale $\gamma$ and shift $\beta$:
$$ \operatorname{LN}(x)_i = \gamma_i ,\frac{x_i - \mu}{\sqrt{\sigma^2 + \epsilon}} + \beta_i, \qquad \mu = \frac1D\sum_i x_i,\quad \sigma^2 = \frac1D\sum_i (x_i-\mu)^2 . $$
For $x = [1, 2, 3]$: $\mu = 2$, $\sigma^2 = 2/3$, and the normalized vector is $[-1.2247, 0, 1.2247]$ (before $\gamma$ and $\beta$). Note that the variance divides by $D$, not $D-1$; the “sample variance” correction from statistics would be a different operation and would break parity with real checkpoints. The small $\epsilon$ (1e-5 in GPT-2) prevents division by zero.
def layernorm(x, gamma, beta, eps=1e-5): # x [..., D]
mean = x.mean(-1, keepdim=True)
var = x.var(-1, keepdim=True, unbiased=False) # divide by D, not D-1
return gamma * (x - mean) / torch.sqrt(var + eps) + beta
inline std::vector<float> layernorm(const std::vector<float>& x, const std::vector<float>& g, const std::vector<float>& b, float eps) {
float mean = std::accumulate(x.begin(), x.end(), 0.f) / x.size(), var = 0;
for (float v : x) var += (v - mean) * (v - mean);
float inv = 1.f / std::sqrt(var / x.size() + eps);
std::vector<float> y(x.size());
for (size_t i = 0; i < x.size(); ++i) y[i] = (x[i] - mean) * inv * g[i] + b[i];
return y;
}
inline std::vector<float> rmsnorm(const std::vector<float>& x, const std::vector<float>& w, float eps) {
float ms = 0;
for (float v : x) ms += v * v;
float inv = 1.f / std::sqrt(ms / x.size() + eps);
std::vector<float> y(x.size());
for (size_t i = 0; i < x.size(); ++i) y[i] = x[i] * inv * w[i];
return y;
}
#![allow(unused)]
fn main() {
/// LayerNorm over one token vector: subtract the mean, divide by the standard deviation.
pub fn layernorm(x: &[f32], gamma: &[f32], beta: &[f32], eps: f32) -> Vec<f32> {
let n = x.len() as f32;
let mean = x.iter().sum::<f32>() / n;
let var = x.iter().map(|v| (v - mean) * (v - mean)).sum::<f32>() / n;
let inv = 1.0 / (var + eps).sqrt();
x.iter().zip(gamma).zip(beta).map(|((v, g), b)| (v - mean) * inv * g + b).collect()
}
/// RMSNorm: no mean subtraction, divide by the root mean square.
pub fn rmsnorm(x: &[f32], weight: &[f32], eps: f32) -> Vec<f32> {
let ms = x.iter().map(|v| v * v).sum::<f32>() / x.len() as f32;
let inv = 1.0 / (ms + eps).sqrt();
x.iter().zip(weight).map(|(v, w)| v * inv * w).collect()
}
}
(The C++ and Rust tabs also include RMSNorm, the simpler variant Qwen uses. It’s Chapter 17’s topic.)
The MLP: where most parameters live
Attention moves information between positions. The MLP (also called the feed-forward network) then processes each position on its own: expand to $4D$, apply a nonlinearity, contract back to $D$:
$$ \operatorname{MLP}(x) = W_{\text{down}},\operatorname{GELU}(W_{\text{up}},x + b_{\text{up}}) + b_{\text{down}} . $$
GPT-2 uses GELU, a smooth relative of ReLU. It’s close to 0 for very negative inputs, close to $x$ for positive ones, and smoothly curved in between. GPT-2 uses the tanh approximation:
$$ \operatorname{GELU}(x) \approx \tfrac12 x\left(1 + \tanh!\left(\sqrt{2/\pi},(x + 0.044715x^3)\right)\right). $$
The exact error-function version differs slightly. Using the wrong one is a small but real source of mismatch when loading checkpoints, so match the checkpoint’s convention (approximate="tanh" for GPT-2).
The MLP holds two thirds of each block’s weights ($8D^2$ versus attention’s $4D^2$). A useful mental model, explored in interpretability research, treats it as a key-value memory: the first matrix detects patterns, and the second writes associated information into the stream.
The block in code
Here is the reference block. Attention packs Q, K and V into one linear layer (qkv, width $3D$), which is a storage choice; GPT-2’s checkpoint stores it that way too. Then come the residual additions you just saw:
class GPTAttention(nn.Module):
"""Multi-head causal self-attention with one packed QKV projection."""
def __init__(self, cfg):
"""Create qkv: width -> 3*width and proj: width -> width, both with bias. (Your engine: Chapter 6)"""
super().__init__()
if cfg.width % cfg.heads:
raise ValueError("width must be divisible by heads")
self.heads = cfg.heads
self.qkv = nn.Linear(cfg.width, 3 * cfg.width)
self.proj = nn.Linear(cfg.width, cfg.width)
def forward(self, x, positions, cache=None, layer=0, rows=None):
"""Project, split heads, (cache), attend, merge heads, project. (Your engine: Chapter 6)"""
q, k, v = self.qkv(x).chunk(3, dim=-1)
q, k, v = (split_heads(t, self.heads) for t in (q, k, v))
key_positions = None
if cache is not None:
k, v, key_positions = cache.update(layer, k, v, positions, rows)
y = causal_attention(q, k, v, positions, key_positions)
return self.proj(merge_heads(y))
class GPTBlock(nn.Module):
def __init__(self, cfg):
"""Two LayerNorms, attention, and a 4x-wide GELU MLP. (Your engine: Chapter 6)"""
super().__init__()
self.ln1 = nn.LayerNorm(cfg.width, eps=cfg.eps)
self.attn = GPTAttention(cfg)
self.ln2 = nn.LayerNorm(cfg.width, eps=cfg.eps)
self.up = nn.Linear(cfg.width, 4 * cfg.width)
self.down = nn.Linear(4 * cfg.width, cfg.width)
self.drop = nn.Dropout(cfg.dropout)
def forward(self, x, positions, cache=None, layer=0, rows=None):
"""Pre-norm residual block: x + attn(ln1(x)), then + mlp(ln2(.)). (Your engine: Chapter 6)"""
x = x + self.drop(self.attn(self.ln1(x), positions, cache, layer, rows))
hidden = nn.functional.gelu(self.up(self.ln2(x)), approximate="tanh")
return x + self.drop(self.down(hidden))
The positions, cache and rows arguments let the same block serve generation with a KV cache later (Chapter 16) and batched serving (Chapter 24). For now they’re just passed through. cache is None and positions are $0..T-1$.
Assembling the model
class GPT(nn.Module):
def __init__(self, cfg):
"""Token and position embeddings, a stack of blocks, a final LayerNorm and a vocabulary head. (Your engine: Chapter 6)"""
super().__init__()
self.cfg = cfg
self.token = nn.Embedding(cfg.vocab, cfg.width)
self.position = nn.Embedding(cfg.context, cfg.width)
self.drop = nn.Dropout(cfg.dropout)
self.blocks = nn.ModuleList(GPTBlock(cfg) for _ in range(cfg.layers))
self.norm = nn.LayerNorm(cfg.width, eps=cfg.eps)
self.head = nn.Linear(cfg.width, cfg.vocab, bias=False)
if cfg.tie_weights:
self.head.weight = self.token.weight
self.apply(self._initialize)
@staticmethod
def _initialize(module):
if isinstance(module, (nn.Linear, nn.Embedding)):
nn.init.normal_(module.weight, std=0.02)
if isinstance(module, nn.Linear) and module.bias is not None:
nn.init.zeros_(module.bias)
def forward(self, ids, cache=None, positions=None, rows=None):
"""ids [B, T] -> logits [B, T, vocab]. (Your engine: Chapter 6)
Positions default to 0..T-1, or continue after the cache's current length.
"""
if positions is None:
start = cache.length if cache is not None else 0
if start + ids.shape[1] > self.cfg.context:
raise ValueError(f"Sequence exceeds the model's context of {self.cfg.context}")
positions = torch.arange(start, start + ids.shape[1], device=ids.device)
# Explicit position tensors are trusted: checking them would force a GPU->CPU sync.
x = self.drop(self.token(ids) + self.position(positions))
for i, block in enumerate(self.blocks):
x = block(x, positions, cache, i, rows)
return self.head(self.norm(x))
@property
def context_limit(self):
return self.cfg.context
def cache_spec(self):
"""(layers, kv_heads, head_dim): what every cache needs to know about this model."""
return self.cfg.layers, self.cfg.heads, self.cfg.width // self.cfg.heads
def new_cache(self, batch, capacity):
if capacity > self.context_limit:
raise ValueError("Requested cache exceeds the context limit")
p = self.token.weight
layers, kv_heads, head_dim = self.cache_spec()
return KVCache(layers, batch, kv_heads, capacity, head_dim, p.device, p.dtype)
Four details deserve attention:
- Two embeddings are added: token identity plus learned position. The position table has
contextrows, which hard-limits the sequence length. - The head produces logits, one score per vocabulary entry at every position: shape
[B, T, V]. Position $t$’s logits predict token $t+1$. - Weight tying:
self.head.weight = self.token.weightmakes the output head the same tensor as the embedding table, read in reverse. It saves $V \cdot D$ parameters (38.6M for GPT-2) and was used by GPT-2. The parameter count must count it once. - Initialization: weights start as small random normals (std 0.02, as in GPT-2) and biases at zero. Getting this wrong produces a model that trains slowly or not at all, but it doesn’t matter for loading pretrained weights, which overwrite everything.
Counting parameters exactly
For a block with biases and two LayerNorms:
| part | parameters |
|---|---|
attention: qkv ($D \times 3D$ + bias) and proj ($D\times D$ + bias) | $4D^2 + 4D$ |
MLP: up ($D \times 4D$ + bias) and down ($4D \times D$ + bias) | $8D^2 + 5D$ |
| two LayerNorms ($\gamma, \beta$ each) | $4D$ |
| block total | $12D^2 + 13D$ |
The whole tied model, with vocabulary $V$ and context $C$:
$$ N = VD + CD + L(12D^2 + 13D) + 2D . $$
For GPT-2 small ($V = 50{,}257$, $C = 1{,}024$, $D = 768$, $L = 12$), that’s 124,439,808, the “124M” in its name. python run.py gpt prints this alongside the count from your model, and your milestone tests that the two agree.
Notice how much of a small model is embedding: 31% of GPT-2 small. As models grow, the $12LD^2$ term dominates, which is why “parameters” and “FLOPs per token” are both close to $12LD^2$ for large models.
A random model already “works”
Before training, your GPT runs end to end and produces logits of the right shape, and the output is noise. That’s useful, not useless: plumbing correctness and learned behavior are separate claims, and you can test the first without the second. The milestone tests check shapes, causality (changing the last token never changes earlier logits), the parameter formula, weight tying, and the identity behavior of zeroed branches. None of that needs training.
Build it
Engine milestone 6: a complete GPT. In engine/gpt.py, implement __init__ and forward for GPTAttention, GPTBlock and GPT. Use your split_heads, merge_heads and causal_attention from Chapter 5. The config dataclass, initialization, parameter-counting helpers and cache factory are provided.
pytest tests/test_ch06_gpt.py
python run.py gpt --impl engine
Tip
Keep the exact attribute names of the reference (
token,position,blocks,norm,head,qkv,proj,ln1,ln2,up,down). Chapter 9’s GPT-2 loader maps checkpoint tensors onto them. If you name things differently, update the loader’s mapping.
Stretch exercises
- ★ Compute by hand the parameter count of GPT-2 medium ($D = 1024$, $L = 24$, 16 heads) and check it with
gpt_parameter_formula. Is it really “355M”? Where: paper, thenexperiments/ch06.py(create it) usingengine.gpt.gpt_parameter_formula. - ★★ Add a
return_hiddenoption that returns the residual stream after every block. Plot each block’s vector norm for a random input. How does it grow with depth? Where:GPT.forwardinengine/gpt.py. - ★★ Convert the block to post-norm (normalize after each addition) and verify that zeroed branches no longer give an identity block. Why? Where:
GPTBlock.forwardinengine/gpt.py; keep a separate post-norm variant for comparison. - ★★★ Implement the exact GELU and measure the maximum difference from the tanh version over [−6, 6]. Then estimate how much it would change GPT-2’s logits. Where:
experiments/ch06.py(create it) for the numerical comparison; change the GELU inGPTBlock.forwardinengine/gpt.pyfor the logit comparison.
Check your understanding
- Which axis does the MLP mix, and which does attention mix?
- Why does a pre-norm block with both branch outputs at zero leave its input unchanged?
- What changes, in storage and in training, when the head is tied to the token embedding?
- Why is LayerNorm’s output not always mean-zero and unit-variance?
- Where do the $12D^2$ parameters of a block come from?
Going deeper
- BALLM Chapter 4 (pp. 92-127): LayerNorm, GELU, shortcut connections and the GPT model, built up in the same order. Its parameter-counting exercise matches this chapter’s formula.
- PMPP §20.1 (pp. 478-482): the decoder block as a sequence of GEMMs and elementwise operations.
- Radford et al., Language Models are Unsupervised Multitask Learners (GPT-2, 2019); He et al., Deep Residual Learning (2015); Ba et al., Layer Normalization (2016).
- Anthropic, A Mathematical Framework for Transformer Circuits (2021): the residual stream view of transformers.