Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

E. Sources, credits and further reading

This book stands on three main sources. Each chapter’s Going deeper section points to the exact sections; this appendix gives the overall map, so you can read the sources alongside the book.

  • BALLM: Sebastian Raschka, Build a Large Language Model (From Scratch), Manning, 2024. The modeling foundations: tokenization, attention, GPT, pretraining, fine-tuning. Companion code: github.com/rasbt/LLMs-from-scratch (Apache-2.0).
  • PMPP: Wen-mei W. Hwu, David B. Kirk and Izzat El Hajj, Programming Massively Parallel Processors: A Hands-on Approach, 5th edition, Morgan Kaufmann. The GPU foundations: CUDA, memory, tiling, reductions, scan, and (Chapter 20) attention and KV caching.
  • GPU Mode: the lecture series and community at github.com/gpu-mode/lectures and youtube.com/@GPUMODE. Kernels, profiling and inference systems, taught by the people who build them.

Chapter map

book chapterBALLMPMPPGPU Mode
1. The big pictureCh. 1§20.4, §20.6L1
2. TensorsApp. A §§A.1-A.3
3. How networks learnApp. A §§A.3-A.7L6
4. Text as numbersCh. 2
5. AttentionCh. 3§20.2
6. The transformer block and GPTCh. 4§20.1
7. TrainingCh. 5 §§5.1-5.2
8. GenerationCh. 5 §5.3
9. Real weightsCh. 5 §§5.4-5.5
10. GPU performanceCh. 1, Ch. 4, Ch. 5 (roofline)L1, L8, L16
11. CUDACh. 2-4L2, L3, L4, L5
12. Fast matrix multiplicationCh. 5-6, Ch. 15L5, L8, L23
13. Reductions and numericsCh. 10, App. AL9
14. Triton§15.8L14, L18, L28, L29
15. FlashAttention§20.5L12, L13, L36
16. KV cache§20.4, §§20.6-20.7
17. Qwen3ch05/11_qwen3 (repository)
18. Engine v1Ch. 5 §5.5§20.6
19. Fast decodeL1, L6, L16, L35
20. QuantizationApp. AL7, L30, L33
21. Fine-tuningCh. 6, Ch. 7
22. LoRA and QLoRAApp. EL32
23. Inside the model
24. Batching§20.6L35
25. Paged attention§20.7L35, L40
26. Speculative decoding§20.6L22
27. Mixture of expertsL11
28. Linear attentionCh. 11 (scan)L20, L21, L24
29. Sparse attention
30. Capstone
31. Engine v2§§20.6-20.7L35, L40
32. Attention backends§20.5L12, L13, L36
33. Fusion, graphs, async schedulingCh. 4, Ch. 15L1, L6, L35
34. Sampling and structured outputCh. 5 §5.3
35. Text boundaryCh. 2
36. HTTP server§20.6L35
37. Batched speculation§20.6L22
38. GGUF and GGMLApp. AL7, L30
39. Quantized checkpointsApp. AL7, L30, L33
40. Hardware and offloadCh. 4, Ch. 5L8, L16
41. Multi-GPUL17
42. Model coverage and MLA
43. Images, pooling and multi-LoRAApp. E (LoRA)L32
44. Benchmarking and hardeningCh. 5 (measurement)L1, L16
45. What’s next

Papers and documentation by topic

Transformers and language models. Vaswani et al., Attention Is All You Need (2017); Radford et al., Language Models are Unsupervised Multitask Learners (GPT-2, 2019); the Qwen3 technical report (2025) and the Qwen3-Next, Qwen3.5 and Qwen3.8-Flash-Next model cards; Su et al., RoFormer (RoPE, 2021); Shazeer, GLU Variants Improve Transformer (2020).

Kernels and performance. Williams, Waterman and Patterson, Roofline (2009); Tillet et al., Triton (2019); Dao et al., FlashAttention 1-3 (2022-2024); Milakov and Gimelshein, Online normalizer calculation for softmax (2018); the CUDA C++ Programming Guide and Best Practices Guide; the Triton tutorials.

Inference systems. Pope et al., Efficiently Scaling Transformer Inference (2022); Yu et al., Orca (2022); Kwon et al., PagedAttention (2023); Zheng et al., SGLang (2024); Agrawal et al., Sarathi-Serve (2024); Zhong et al., DistServe (2024); Leviathan et al. and Chen et al. on speculative decoding (2023); the PyTorch gpt-fast blog posts.

Quantization and adaptation. Dettmers et al., LLM.int8() (2022) and QLoRA (2023); Frantar et al., GPTQ (2022); Lin et al., AWQ (2023); Xiao et al., SmoothQuant (2022); Hu et al., LoRA (2021).

Architectures. Shazeer et al. (2017) and Fedus et al. (2021) on mixtures of experts; DeepSeek-V3 technical report; Katharopoulos et al. (2020), Schlag et al. (2021), Yang et al. (2023-2024) on linear attention and Gated DeltaNet; Gu and Dao on Mamba (2023-2024); Zhu et al., Hyper-Connections (2024); DeepSeek’s Native Sparse Attention and Engram (2025-2026).

Interpretability. Elhage et al., A Mathematical Framework for Transformer Circuits (2021); nostalgebraist, the logit lens (2020); Turner et al., Activation Addition (2023); Arditi et al., Refusal Is Mediated by a Single Direction (2024).

Credits

  • The Verdict by Edith Wharton (1908), public domain, used as BALLM uses it for the training chapters.
  • The 1,100-example instruction dataset (data/instruction-data.json) comes from BALLM Chapter 7’s companion repository (Apache-2.0).
  • The SMS Spam Collection (Almeida and Hidalgo, UCI Machine Learning Repository, CC BY 4.0) is referenced for Chapter 21 and downloaded by the reader.
  • Kernel structure follows the official Triton tutorials (fused softmax, matmul with grouped ordering, fused attention) and PMPP’s CUDA examples, adapted and simplified.
  • The Flash-Next implementation was written from the Transformers qwen4_exp reference implementation (Apache-2.0) and is tested against it.
  • Diagrams and interactive explorers are original to this book.

Part VIII primary sources

The foundations above remain useful, but Part VIII also follows the code and papers of the systems it compares. Chapter-level links give the particular mechanism; the following groups are the starting points.