Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

23. Inside the model: hooks, the logit lens and steering

In this chapter

  • The residual stream as a shared workspace that every layer reads and writes, and how to observe it with forward hooks.
  • The logit lens: decoding what the model "would say" at every layer.
  • Directions in activation space: finding one from contrasting examples, adding it to steer behavior, projecting it out to remove behavior.
  • Making an edit permanent by orthogonalizing weights, and evaluating edits honestly.

You will build

direction_from_means and project_out in engine/steering.py. You'll read your model's layers with the logit lens, double the probability of an all-caps answer with one added vector, and remove a direction from a real Qwen3.

Time: 4-5 hours. GPU: optional.

The residual stream

Look at Chapter 6’s block again: x = x + attention(norm(x)), then x = x + mlp(norm(x)). Every layer reads the current $x$ and adds its result back. Nothing ever overwrites $x$. So the hidden state after layer $L$ is a sum:

$$ x_L = \underbrace{x_0}{\text{embedding}} + \sum{\ell=1}^{L} \big(\text{attn}\ell + \text{mlp}\ell\big). $$

This residual stream is a shared workspace: a $d$-dimensional vector per position that each component reads from and writes to, and that the final norm and head decode into the next-token prediction. Two consequences make this chapter possible:

  1. Because the head reads the stream through one linear map, you can apply it to the stream at any layer and see what’s there. That’s the logit lens.
  2. Because components communicate by adding vectors, a feature the model uses is often represented as a direction in this space (the linear representation hypothesis). Adding or removing a direction can change behavior directly.

Hooks: observing without changing code

PyTorch’s forward hooks run a function after a module’s forward pass, with access to its inputs and output. If the function returns a value, it replaces the output. That’s all you need to read and edit activations in any model, yours or a library’s, without changing its source:

@contextmanager
def capture(module, store, name):
    """Record a module's output into store[name] for every forward call inside the block."""
    def hook(_module, _inputs, output):
        store.setdefault(name, []).append((output[0] if isinstance(output, tuple) else output).detach())
    handle = module.register_forward_hook(hook)
    try:
        yield store
    finally:
        handle.remove()                       # always detach, even when the forward raises


@torch.no_grad()
def residuals(model, ids, layers=None):
    """Hidden state after each requested decoder layer, as {layer_index: [B, T, D]}."""
    blocks = list(_blocks(model))
    layers = range(len(blocks)) if layers is None else layers
    store, handles = {}, []
    for i in layers:
        handles.append(blocks[i].register_forward_hook(
            lambda _m, _inp, out, i=i: store.__setitem__(i, (out[0] if isinstance(out, tuple) else out).detach())))
    try:
        model(ids)
    finally:
        for handle in handles:
            handle.remove()
    return store

Two habits matter. Always remove hooks in a finally block: a forgotten hook keeps editing every later forward pass, and an exception inside a forward is exactly when you’d forget. And detach() anything you store, or you’ll keep the whole autograd graph alive.

The logit lens

Apply the final norm and the head to the residual stream after each layer, and look at the top tokens:

@torch.no_grad()
def logit_lens(model, ids, top=5):
    """Decode every layer's residual through the final norm and head: what would the model
    predict if it stopped here? Returns {layer: top token IDs at the last position}."""
    norm = model.norm if hasattr(model, "norm") else model.model.norm
    head = model.head if hasattr(model, "head") else model.lm_head
    return {layer: head(norm(h[:, -1])).topk(top, dim=-1).indices[0].tolist()
            for layer, h in residuals(model, ids).items()}

On a real model the picture is striking (nostalgebraist, 2020). Early layers predict tokens related to the current input; middle layers converge toward plausible continuations; the last few layers sharpen the final answer. For GPT-2 on “The Eiffel Tower is in the city of”, “ Paris“ typically rises to the top only in the later layers. Different models use their early layers differently, and the plain logit lens can be misleading for some; the tuned lens (Belrose et al., 2023) learns a small linear translator per layer to fix this.

Your 4-layer instruction-tuned GPT from Chapter 21, teacher-forced through “The contraction for”:

python run.py steer
{"text_end": " Response:\nThe contraction for", "layer_0": [" '", " \"", "'"], "layer_1": [" '", " \"", "'"],
 "layer_2": [" '", " \"", " the"], "layer_3": [" '", " \"", " the"]}

Every layer already expects a quote, the dataset’s style for this kind of answer (“The contraction for ‘it is’ is …”). A small model decides early. The last layer’s lens equals the real prediction exactly, which the milestone test checks.

Directions from contrasting examples

How do you find the direction for a concept? The simplest method that works surprisingly well is a difference of means. Collect activations at one layer for examples with the property (positive) and without it (negative), and take the normalized difference of their averages:

$$ d = \frac{\bar{h}\text{pos} - \bar{h}\text{neg}}{\lVert \bar{h}\text{pos} - \bar{h}\text{neg} \rVert}. $$

The pairs should differ only in the property, so everything else averages out: “Explain recursion in a cheerful conversational tone” versus “… in a formal technical tone”, for 32 topics (data/style-positive.jsonl and style-negative.jsonl).

Two edits use a direction:

  • Steering (adding): $h \leftarrow h + \lambda d$ at one layer, at every position. Positive $\lambda$ pushes the model toward the property.
  • Directional ablation (projecting out): $h \leftarrow h - (h \cdot d), d$. The component along $d$ becomes exactly zero, so downstream layers can no longer read the property from it.
def direction_from_means(positive, negative):
    """Unit vector from the mean of negative examples toward the mean of positive ones.  (Your engine: Chapter 23)"""
    delta = positive.float().mean(0) - negative.float().mean(0)
    norm = delta.norm()
    if not torch.isfinite(norm) or norm < 1e-8:
        raise ValueError("The two groups have (almost) the same mean: no usable direction")
    return delta / norm


def project_out(hidden, direction, strength=1.0):
    """h - strength * (h . d) d. With strength 1 the d-component becomes exactly zero.  (Your engine: Chapter 23)"""
    h = hidden.float()
    return (h - strength * (h @ direction.float()).unsqueeze(-1) * direction.float()).to(hidden.dtype)
@contextmanager
def steer(model, layer, direction, mode="add", strength=4.0):
    """Temporarily edit one block's output: 'add' pushes along the direction, 'ablate' removes it."""
    block = _blocks(model)[layer]

    def hook(_module, _inputs, output):
        hidden = output[0] if isinstance(output, tuple) else output
        if mode == "add":
            edited = hidden + strength * direction.to(hidden)
        elif mode == "ablate":
            edited = project_out(hidden, direction.to(hidden.device), strength)
        else:
            raise ValueError("mode must be 'add' or 'ablate'")
        return (edited, *output[1:]) if isinstance(output, tuple) else edited
    handle = block.register_forward_hook(hook)
    try:
        yield model
    finally:
        handle.remove()

Explore: a direction in activation space

Two clouds of 2-D activations, positive and negative. See the difference-of-means direction, then drag the steering strength or switch to ablation and watch every point move.

Run it

run.py steer builds a direction at layer 2 of your Chapter 21 model from 300 training answers, written normally (negative) and in capitals (positive). Then it adds the direction, scaled by multiples of the gap between the two means, and measures the probability that the first response token is in capitals, over 60 test prompts:

{"layer": 2, "added": "0.0 x gap (0.0)",  "first_token_all_caps_probability": 0.2227}
{"layer": 2, "added": "0.5 x gap (5.4)",  "first_token_all_caps_probability": 0.2877}
{"layer": 2, "added": "1.0 x gap (10.8)", "first_token_all_caps_probability": 0.3699}
{"layer": 2, "added": "1.5 x gap (16.1)", "first_token_all_caps_probability": 0.4431}

One vector, no training, and the probability doubles. (The baseline is 22% because single capital-letter tokens like “A” count.) In a model this small, though, the direction is crude: push harder, or generate long texts, and fluency collapses before the style changes cleanly. Steering works far better in large models, whose representations are more linear and more robust. Turner et al. (2023) steered GPT-2 XL’s topic and sentiment with a single prompt-pair difference, and Arditi et al. (2024) found that refusal in 13 open chat models is mediated by one direction.

Making an edit permanent

A hook works at runtime. To ship an edit, change the weights. Every write into the residual stream comes from a matrix: the embedding, each attention output projection and each MLP down projection. If each writer’s output has no component along $d$, the stream never gains one. For a linear layer $y = Wx + b$, replace

$$ W \leftarrow (I - d d^\top) W, \qquad b \leftarrow (I - d d^\top), b , $$

so that $d^\top y = 0$ for every input:

@torch.no_grad()
def orthogonalize_output(linear, direction):
    """Make a linear layer unable to write along `direction`: W <- (I - d d^T) W, b <- (I - d d^T) b.
    Applied to every module that writes into the residual stream, this makes ablation permanent."""
    d = direction.to(linear.weight)
    linear.weight -= torch.outer(d, d @ linear.weight)
    if linear.bias is not None:
        linear.bias -= d * (d @ linear.bias)

Applied to all writers (and the embedding’s rows), this is weight orthogonalization, popularized as “abliteration” after Arditi et al.’s refusal paper. The milestone test checks that the edited layer equals the runtime projection exactly.

Warning

Directional ablation of a refusal direction removes a model’s safety behavior. It’s a well-documented interpretability result, and it’s why open-weight safety can’t rely on refusals alone. This book uses tone and capitalization as its examples; if you study refusal, do it on models and in settings where removing safeguards is appropriate, and don’t distribute such edits.

Editing a real Qwen3

edit_checkpoint.py runs the whole procedure on a Hugging Face Qwen3 checkpoint: it fits a direction at one block’s output from the last prompt token of the positive and negative prompts, ablates it with a hook on every position, and compares the edited and original models on held-out data:

python edit_checkpoint.py --model-dir models/Qwen3-0.6B --layer 14 \
    --positive-file data/style-positive.jsonl --negative-file data/style-negative.jsonl \
    --valid-file data/instruction-valid.jsonl --output runs/tone-direction.pt

It reports response loss before and after, the KL divergence from the original model on held-out responses, one greedy generation each way, and saves the direction with a manifest. No weights are changed; apply orthogonalize_output to make it permanent once you’re convinced.

Evaluate edits like any other change

  • Did it do what you wanted? Measure the targeted behavior on held-out prompts, not the ones used to fit the direction.
  • What else did it change? Report KL from the original on unrelated data, and held-out loss. A direction that changes everything is not a “tone” direction.
  • Which layer? Directions work best in the middle layers; sweep a few and pick by held-out measurements, not by eye.
  • Combine carefully. Ablation, LoRA (Chapter 22) and quantization (Chapter 20) don’t commute. Evaluate the artifact you ship.

Build it

Engine milestone 23: activation editing. Implement direction_from_means and project_out in engine/steering.py (hooks, the logit lens and weight orthogonalization are provided).

pytest tests/test_ch23_steering.py
python run.py steer --impl engine

The tests check that projecting out removes the component exactly and is idempotent, that hooks are removed even when the forward raises, that orthogonalized weights equal the runtime projection, and that the logit lens at the last layer equals the model’s real prediction.

Stretch exercises

  1. ★ Run run.py steer --layer L for every layer, and plot the all-caps probability at 1.0 × gap against the layer. Where: terminal: python run.py steer --impl engine --layer L; plot results in experiments/ch23.py (create it).
  2. ★★ Load GPT-2 (Chapter 9) and print the logit lens for “The Eiffel Tower is in the city of”. At which layer does “ Paris“ enter the top 5? Where: experiments/ch23.py (create it), calling engine.loaders.load_gpt2 and engine.steering.logit_lens.
  3. ★★ Steer Qwen3-0.6B’s tone with the style prompts: add $\lambda d$ at the best layer during generation and compare five answers at $\lambda = 0$, 4 and 8. Where: experiments/ch23.py (create it), using engine.steering.steer with engine.engine.LLM.
  4. ★★★ Orthogonalize every residual writer of your Qwen3 engine against a direction, save the edited checkpoint, and verify its logits equal the hooked model’s. Where: extend orthogonalize_output in engine/steering.py for Qwen3’s residual writers; save with engine.safetensors_io.save_file.

Check your understanding

  1. Why can the final norm and head be applied to the residual stream after any layer?
  2. Why should positive and negative examples differ only in the property you want?
  3. What’s the difference between adding a direction and projecting it out?
  4. Why does orthogonalizing every residual writer make ablation permanent?
  5. How would you show that an edit changed only what you intended?

Going deeper

  • nostalgebraist, interpreting GPT: the logit lens (2020); Belrose et al., Eliciting Latent Predictions from Transformers with the Tuned Lens (2023).
  • Elhage et al., A Mathematical Framework for Transformer Circuits (Anthropic, 2021): the residual-stream view this chapter uses.
  • Turner et al., Activation Addition: Steering Language Models Without Optimization (2023); Zou et al., Representation Engineering (2023); Arditi et al., Refusal in Language Models Is Mediated by a Single Direction (2024).
  • Sparse autoencoders, which find thousands of interpretable directions at once: Bricken et al., Towards Monosemanticity (2023) and Templeton et al., Scaling Monosemanticity (2024); Gemma Scope and the TransformerLens and nnsight libraries for doing this on real models.