Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

22. LoRA and QLoRA

In this chapter

  • Why fine-tuning updates can be low-rank, and how LoRA trains a model through two small matrices per layer.
  • Initialization, the α/r scale, and which layers to adapt.
  • Merging an adapter into the weights for zero inference cost, or keeping it separate to serve many adapters on one base.
  • QLoRA: training adapters on a 4-bit frozen base, and the memory arithmetic of all three approaches.

You will build

LoRALinear in engine/lora.py: the adapted forward pass and the merge. You'll teach your instruction-tuned model a new style with 5% of its parameters, merge the adapter and verify that nothing changed.

Time: 3-5 hours. GPU: recommended for Qwen3 (the demo runs on a CPU in under a minute).

The cost of full fine-tuning

Chapter 21 ended with an uncomfortable table: full fine-tuning with AdamW needs about 18 bytes per parameter, 144 GB for an 8B model before activations. And every fine-tuned variant is a complete copy of the model: ten customers with ten fine-tunes of an 8B model means 160 GB of checkpoints.

LoRA (Low-Rank Adaptation, Hu et al., 2021) fixes both. It freezes every pretrained weight and learns a small correction for selected matrices. Only the corrections get gradients and optimizer state, and only they are saved.

The low-rank idea

Fine-tuning changes a weight matrix $W \in \mathbb{R}^{d_\text{out} \times d_\text{in}}$ to $W + \Delta W$. Empirically, the useful $\Delta W$ for adapting a pretrained model has low intrinsic rank (Aghajanyan et al., 2020): it can be well approximated by a product of two thin matrices. So LoRA parameterizes it that way:

$$ W’ = W + \frac{\alpha}{r} B A, \qquad A \in \mathbb{R}^{r \times d_\text{in}},\ B \in \mathbb{R}^{d_\text{out} \times r},\ r \ll \min(d_\text{in}, d_\text{out}). $$

For a 4,096 × 4,096 projection, full fine-tuning trains 16.8M numbers; LoRA with $r = 8$ trains $8 \times (4096 + 4096) = 65{,}536$, which is 0.39%.

The forward pass never forms $BA$. It computes the frozen path and the low-rank path separately and adds them:

$$ y = W x + \frac{\alpha}{r}, B (A x). $$

$Ax$ is an $r$-vector, cheap to compute. The extra FLOPs are $2r(d_\text{in} + d_\text{out})$ per token, under 1% of the layer.

class LoRALinear(nn.Module):
    """y = base(x) + (alpha / r) * B(A(x)), with the base frozen.  (Your engine: Chapter 22)

    A [r, in] starts random and B [out, r] starts at zero, so the adapted layer initially
    equals the base exactly; B receives a gradient on the first step, A on the next.
    """

    def __init__(self, base, rank=8, alpha=16):
        super().__init__()
        if rank < 1:
            raise ValueError("rank must be positive")
        self.base = base
        for parameter in base.parameters():
            parameter.requires_grad_(False)
        self.scale = alpha / rank
        self.A = nn.Parameter(base.weight.new_empty(rank, base.in_features))
        self.B = nn.Parameter(base.weight.new_zeros(base.out_features, rank))
        nn.init.kaiming_uniform_(self.A, a=math.sqrt(5))

    def forward(self, x):
        """(Your engine: Chapter 22)"""
        return self.base(x) + F.linear(F.linear(x, self.A), self.B) * self.scale

    @torch.no_grad()
    def merged(self):
        """A plain nn.Linear with W + scale * B @ A folded in. Do not also keep the adapter active.  (Your engine: Chapter 22)"""
        out = nn.Linear(self.base.in_features, self.base.out_features, bias=self.base.bias is not None,
                        device=self.base.weight.device, dtype=self.base.weight.dtype)
        out.weight.copy_(self.base.weight + self.scale * (self.B @ self.A))
        if self.base.bias is not None:
            out.bias.copy_(self.base.bias)
        return out

Initialization and the scale

  • $B = 0$, $A$ random. The adapted model starts exactly equal to the base, so training begins from the pretrained behavior. On the first step, $B$ gets a gradient ($\partial L / \partial B = \frac{\alpha}{r} g, (Ax)^\top$, which is non-zero) while $A$’s gradient is zero ($\partial L/\partial A = \frac{\alpha}{r} B^\top g, x^\top$, and $B = 0$). After one update, $B \ne 0$ and both train. The milestone test checks this.
  • $\alpha/r$. Scaling the update by $\alpha/r$ keeps its size roughly stable when you change $r$, so a learning rate tuned at one rank still works at another. A common choice is $\alpha = 2r$. (rsLoRA, Kalajdzievski 2023, argues for $\alpha/\sqrt{r}$ at large ranks.)

Which layers to adapt

def add_lora(model, targets=("q_proj", "v_proj"), rank=8, alpha=16):
    """Freeze the model, then wrap every Linear whose attribute name is in targets."""
    for parameter in model.parameters():
        parameter.requires_grad_(False)
    wrapped = []
    for name, module in list(model.named_modules()):
        for child_name, child in list(module.named_children()):
            if child_name in targets and isinstance(child, nn.Linear):
                setattr(module, child_name, LoRALinear(child, rank, alpha))
                wrapped.append(f"{name}.{child_name}")
    return wrapped


def merge_lora(model):
    """Replace every LoRALinear with its merged Linear, in place."""
    for module in list(model.modules()):
        for child_name, child in list(module.named_children()):
            if isinstance(child, LoRALinear):
                setattr(module, child_name, child.merged())
    return model

The original paper adapted only the attention query and value projections. The QLoRA paper found that adapting all linear layers (attention and MLP) matters more than the rank: with all layers adapted, ranks from 8 to 64 performed similarly. Embeddings and the LM head are usually left alone; for a model with tied embeddings, adapting the head would also change the input embedding through the shared tensor.

Explore: what LoRA trains

Pick a model shape, the adapted layers and the rank. See trainable parameters, adapter file size, and training memory compared with full fine-tuning.

Run it: a new style with 5% of the parameters

Chapter 21’s run.py sft saved its instruction-tuned GPT to runs/sft.pt. Now teach it a new behavior, answering in capital letters, by training LoRA adapters on every linear layer with the same instructions and upper-cased responses:

python run.py lora --steps 300

On a laptop CPU, in 33 seconds:

{"before_lora": "The gas' is 'bua'.", "valid_loss_on_caps": 6.5372, "dropped_overlong": 0}
{"wrapped_layers": 16, "trainable": 98304, "total": 2025600, "trainable_fraction": 0.0485}
{"step": 1, "train_loss": 6.5902, "valid_loss": 6.4817}
{"step": 151, "train_loss": 3.6265, "valid_loss": 3.6524}
{"step": 300, "train_loss": 3.2451, "valid_loss": 3.5377}
{"instruction": "What is the contraction for 'it is'?", "after_lora": "THE OMES 'TES 'S 'S 'TETES 'TETETETETESTETESSTESTESTE"}
{"instruction": "Convert the number 5 from decimal to binary.", "after_lora": "15 MES: 14"}
{"merged_max_abs_difference": 4.410743713378906e-06, "parameters_after_merge": 2025600}

The style moved completely; the content was nonsense before and remains so (this model never pretrained), and the first answer degenerates into repetition, Chapter 8’s classic greedy failure. The fraction is high here only because the model is tiny. For Qwen3-0.6B, rank 8 on q_proj and v_proj trains 1.15M parameters, 0.19% of the model. On pretrained GPT-2 small, BALLM Appendix E’s LoRA (rank 16 on every linear layer, 2.7M trainable parameters, 2% of the model) reaches 98.0% test accuracy on spam classification, slightly above full fine-tuning of the last block in Chapter 6.

The last line is the merge check, discussed next.

Merge, or keep adapters separate

After training, an adapter can be merged: compute $W + \frac{\alpha}{r}BA$ once and replace $W$. The merged model is an ordinary model with zero extra inference cost. The demo’s merge changed the logits by $4 \times 10^{-6}$, floating-point rounding from adding in a different order.

Or keep adapters separate. One base model in GPU memory can then serve hundreds of fine-tunes, each a few megabytes, applying each request’s adapter in its forward pass. Batching requests with different adapters needs a gathered low-rank matmul (S-LoRA and Punica’s BGMV kernels; vLLM and SGLang support this). That’s how serving providers host many customer fine-tunes cheaply.

Rules that prevent subtle bugs:

  • Never apply an adapter to a model it has already been merged into: the update would count twice.
  • A merged model that you then quantize is a different model from the adapter-on-quantized-base you trained with QLoRA. Evaluate the artifact you’ll ship.
  • An adapter is only valid for the exact base weights it was trained on. Record the base’s revision with the adapter (finetune.py writes it to its manifest).

QLoRA: a 4-bit frozen base

LoRA removes gradients and optimizer state for the base, but the base weights themselves still sit in BF16. QLoRA (Dettmers et al., 2023) stores the frozen base in 4 bits and dequantizes on the fly in the forward and backward passes, while the adapters train in BF16. It combines three ideas:

  • NF4, a 4-bit data type whose 16 levels are quantiles of a normal distribution, matching how pretrained weights are distributed (Chapter 20). Unlike Chapter 20’s uniform INT4 grid, dequantizing is a table lookup.
  • Double quantization: the per-block scales (one per 64 weights) are themselves quantized to 8 bits, saving about 0.37 bits per parameter.
  • Paged optimizers: optimizer state that can spill to CPU memory during memory spikes.

Memory per parameter of the base, with activations and the small adapters aside:

methodbase weightsgradients + optimizertotal per base parameter8B model
full fine-tuning (mixed precision)2 + 4 (master)4 + 818 bytes~144 GB
LoRA, BF16 base2~0~2 bytes~16 GB
QLoRA, NF4 base~0.52~0~0.52 bytes~4.2 GB

That’s why QLoRA fine-tunes an 8B model on a single consumer GPU, and why the QLoRA paper could fine-tune a 65B model on one 48 GB GPU. The cost is speed: dequantizing every weight in every forward and backward pass makes each step slower than BF16 LoRA.

LoRA and QLoRA on Qwen3

finetune.py uses Hugging Face PEFT for the adapters, with the same data and loss handling as Chapter 21:

uv pip install -r workflow-requirements.txt
python finetune.py train --mode lora --rank 8 --model-dir models/Qwen3-0.6B \
    --train-file data/instruction-train.jsonl --valid-file data/instruction-valid.jsonl \
    --output runs/qwen3-lora --steps 200 --lr 2e-4 --batch-size 4 --accumulation 4
python finetune.py evaluate --model-dir models/Qwen3-0.6B \
    --valid-file data/instruction-valid.jsonl --output runs/qwen3-lora

The saved directory holds only the adapter (a few MB) and the tokenizer. Note the learning rate: LoRA typically needs 10-20× the learning rate of full fine-tuning, because the update starts at zero and lives in a small subspace.

For QLoRA, install bitsandbytes and pass --mode qlora. The script loads the base with NF4, double quantization and BF16 compute, prepares it for k-bit training (casting norms to FP32, enabling gradient checkpointing), and adds the same adapters.

When bitsandbytes’ CUDA build doesn’t match PyTorch’s

bitsandbytes ships compiled CUDA libraries for specific CUDA versions. If PyTorch uses a newer CUDA than any bundled library, bitsandbytes looks for a file that doesn’t exist (for example libbitsandbytes_cuda132.so with PyTorch on CUDA 13.2) and fails or falls back. Diagnose first:

python -m bitsandbytes            # prints the CUDA version it detected and the library it loaded

If a library for an older compatible CUDA is bundled and that toolkit is installed, the documented override selects it, as in this example for a CUDA 13.0 build:

export BNB_CUDA_VERSION=130
export LD_LIBRARY_PATH="/usr/local/cuda-13.0/lib64${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"
python -m bitsandbytes            # must now report success before you train

This is an environment-specific workaround, not a fix for an unsupported GPU or platform (check the installation guide for ARM64 and newer architectures). If QLoRA can’t run, use plain LoRA; don’t call the result QLoRA.

Build it

Engine milestone 22: LoRA. Implement LoRALinear.forward and LoRALinear.merged in engine/lora.py (the constructor, add_lora and merge_lora are provided).

pytest tests/test_ch22_lora.py
python run.py lora --impl engine --steps 300

The tests check that an adapted layer starts equal to its base, that only $A$ and $B$ receive gradients (and that $A$’s is zero on the first step), that the merged layer matches the live adapter, and that adding and merging adapters on a whole Qwen3 model preserves its outputs.

Stretch exercises

  1. ★ Run run.py lora with ranks 1, 2, 8 and 32, and plot final validation loss against trainable parameters. Where does a larger rank stop helping? Where: terminal: python run.py lora --impl engine --rank R for each rank; record results in experiments/ch22.py (create it).
  2. ★★ Compare full SFT (Chapter 21) and LoRA on GPT-2 with the same data and number of steps: held-out loss, peak memory (torch.cuda.max_memory_allocated) and seconds per step. Where: experiments/ch22.py (create it), adapting run.py’s cmd_sft and cmd_lora.
  3. ★★ Serve two adapters at once: for a batch where row 0 uses adapter 1 and row 1 uses adapter 2, compute $B_{i}(A_{i} x)$ per row with gathered weights, and verify against running each row separately. Where: add a gathered-adapter helper in engine/lora.py; Chapter 43 integrates it in engine/serve/adapters.py.
  4. ★★★ Implement DoRA (Liu et al., 2024): decompose $W$ into magnitude and direction, $W’ = m \cdot \frac{W + BA}{\lVert W + BA \rVert_c}$ with a trainable per-column magnitude $m$, and compare with LoRA at equal parameter count. Where: add a DoRA layer beside LoRALinear in engine/lora.py.

Check your understanding

  1. Why does initializing $B$ to zero make the adapted model equal to the base at the start?
  2. Why is $A$’s gradient zero on the first step, and why doesn’t that stop training?
  3. What does merging cost at inference time, and what does keeping adapters separate buy you?
  4. Why does QLoRA reduce memory but slow each step down?
  5. Why can’t an adapter trained on one base model be applied to another?

Going deeper

  • BALLM Appendix E (Parameter-efficient fine-tuning with LoRA): LoRA from scratch on GPT-2 for spam classification, the source of this chapter’s comparison.
  • Hu et al., LoRA (2021); Dettmers et al., QLoRA (2023); Aghajanyan et al., Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning (2020); Liu et al., DoRA (2024); Sheng et al., S-LoRA (2023) and Chen et al., Punica (2023) for multi-adapter serving.
  • GPU Mode L32 (Unsloth: LLM systems engineering): fused Triton kernels for LoRA and QLoRA training.
  • The Hugging Face PEFT documentation (LoRA configuration, merging, quantized bases) and Sebastian Raschka’s article Practical Tips for Finetuning LLMs Using LoRA.