E. Sources, credits and further reading
This book stands on three main sources. Each chapter’s Going deeper section points to the exact sections; this appendix gives the overall map, so you can read the sources alongside the book.
- BALLM: Sebastian Raschka, Build a Large Language Model (From Scratch), Manning, 2024. The modeling foundations: tokenization, attention, GPT, pretraining, fine-tuning. Companion code: github.com/rasbt/LLMs-from-scratch (Apache-2.0).
- PMPP: Wen-mei W. Hwu, David B. Kirk and Izzat El Hajj, Programming Massively Parallel Processors: A Hands-on Approach, 5th edition, Morgan Kaufmann. The GPU foundations: CUDA, memory, tiling, reductions, scan, and (Chapter 20) attention and KV caching.
- GPU Mode: the lecture series and community at github.com/gpu-mode/lectures and youtube.com/@GPUMODE. Kernels, profiling and inference systems, taught by the people who build them.
Chapter map
| book chapter | BALLM | PMPP | GPU Mode |
|---|---|---|---|
| 1. The big picture | Ch. 1 | §20.4, §20.6 | L1 |
| 2. Tensors | App. A §§A.1-A.3 | ||
| 3. How networks learn | App. A §§A.3-A.7 | L6 | |
| 4. Text as numbers | Ch. 2 | ||
| 5. Attention | Ch. 3 | §20.2 | |
| 6. The transformer block and GPT | Ch. 4 | §20.1 | |
| 7. Training | Ch. 5 §§5.1-5.2 | ||
| 8. Generation | Ch. 5 §5.3 | ||
| 9. Real weights | Ch. 5 §§5.4-5.5 | ||
| 10. GPU performance | Ch. 1, Ch. 4, Ch. 5 (roofline) | L1, L8, L16 | |
| 11. CUDA | Ch. 2-4 | L2, L3, L4, L5 | |
| 12. Fast matrix multiplication | Ch. 5-6, Ch. 15 | L5, L8, L23 | |
| 13. Reductions and numerics | Ch. 10, App. A | L9 | |
| 14. Triton | §15.8 | L14, L18, L28, L29 | |
| 15. FlashAttention | §20.5 | L12, L13, L36 | |
| 16. KV cache | §20.4, §§20.6-20.7 | ||
| 17. Qwen3 | ch05/11_qwen3 (repository) | ||
| 18. Engine v1 | Ch. 5 §5.5 | §20.6 | |
| 19. Fast decode | L1, L6, L16, L35 | ||
| 20. Quantization | App. A | L7, L30, L33 | |
| 21. Fine-tuning | Ch. 6, Ch. 7 | ||
| 22. LoRA and QLoRA | App. E | L32 | |
| 23. Inside the model | |||
| 24. Batching | §20.6 | L35 | |
| 25. Paged attention | §20.7 | L35, L40 | |
| 26. Speculative decoding | §20.6 | L22 | |
| 27. Mixture of experts | L11 | ||
| 28. Linear attention | Ch. 11 (scan) | L20, L21, L24 | |
| 29. Sparse attention | |||
| 30. Capstone | |||
| 31. Engine v2 | §§20.6-20.7 | L35, L40 | |
| 32. Attention backends | §20.5 | L12, L13, L36 | |
| 33. Fusion, graphs, async scheduling | Ch. 4, Ch. 15 | L1, L6, L35 | |
| 34. Sampling and structured output | Ch. 5 §5.3 | ||
| 35. Text boundary | Ch. 2 | ||
| 36. HTTP server | §20.6 | L35 | |
| 37. Batched speculation | §20.6 | L22 | |
| 38. GGUF and GGML | App. A | L7, L30 | |
| 39. Quantized checkpoints | App. A | L7, L30, L33 | |
| 40. Hardware and offload | Ch. 4, Ch. 5 | L8, L16 | |
| 41. Multi-GPU | L17 | ||
| 42. Model coverage and MLA | |||
| 43. Images, pooling and multi-LoRA | App. E (LoRA) | L32 | |
| 44. Benchmarking and hardening | Ch. 5 (measurement) | L1, L16 | |
| 45. What’s next |
Papers and documentation by topic
Transformers and language models. Vaswani et al., Attention Is All You Need (2017); Radford et al., Language Models are Unsupervised Multitask Learners (GPT-2, 2019); the Qwen3 technical report (2025) and the Qwen3-Next, Qwen3.5 and Qwen3.8-Flash-Next model cards; Su et al., RoFormer (RoPE, 2021); Shazeer, GLU Variants Improve Transformer (2020).
Kernels and performance. Williams, Waterman and Patterson, Roofline (2009); Tillet et al., Triton (2019); Dao et al., FlashAttention 1-3 (2022-2024); Milakov and Gimelshein, Online normalizer calculation for softmax (2018); the CUDA C++ Programming Guide and Best Practices Guide; the Triton tutorials.
Inference systems. Pope et al., Efficiently Scaling Transformer Inference (2022); Yu et al., Orca (2022); Kwon et al., PagedAttention (2023); Zheng et al., SGLang (2024); Agrawal et al., Sarathi-Serve (2024); Zhong et al., DistServe (2024); Leviathan et al. and Chen et al. on speculative decoding (2023); the PyTorch gpt-fast blog posts.
Quantization and adaptation. Dettmers et al., LLM.int8() (2022) and QLoRA (2023); Frantar et al., GPTQ (2022); Lin et al., AWQ (2023); Xiao et al., SmoothQuant (2022); Hu et al., LoRA (2021).
Architectures. Shazeer et al. (2017) and Fedus et al. (2021) on mixtures of experts; DeepSeek-V3 technical report; Katharopoulos et al. (2020), Schlag et al. (2021), Yang et al. (2023-2024) on linear attention and Gated DeltaNet; Gu and Dao on Mamba (2023-2024); Zhu et al., Hyper-Connections (2024); DeepSeek’s Native Sparse Attention and Engram (2025-2026).
Interpretability. Elhage et al., A Mathematical Framework for Transformer Circuits (2021); nostalgebraist, the logit lens (2020); Turner et al., Activation Addition (2023); Arditi et al., Refusal Is Mediated by a Single Direction (2024).
Credits
- The Verdict by Edith Wharton (1908), public domain, used as BALLM uses it for the training chapters.
- The 1,100-example instruction dataset (
data/instruction-data.json) comes from BALLM Chapter 7’s companion repository (Apache-2.0). - The SMS Spam Collection (Almeida and Hidalgo, UCI Machine Learning Repository, CC BY 4.0) is referenced for Chapter 21 and downloaded by the reader.
- Kernel structure follows the official Triton tutorials (fused softmax, matmul with grouped ordering, fused attention) and PMPP’s CUDA examples, adapted and simplified.
- The Flash-Next implementation was written from the Transformers
qwen4_expreference implementation (Apache-2.0) and is tested against it. - Diagrams and interactive explorers are original to this book.
Part VIII primary sources
The foundations above remain useful, but Part VIII also follows the code and papers of the systems it compares. Chapter-level links give the particular mechanism; the following groups are the starting points.
- Scheduling and serving: vLLM, SGLang, PagedAttention, Sarathi-Serve and DistServe. Chapters 31-37 implement the shared mechanisms with smaller interfaces.
- Attention: FlashAttention, FlashInfer and FlashMLA. Their architecture-specific kernels are comparison/reference targets, not claims that the teaching kernels reach the same performance.
- Formats and hardware: llama.cpp/GGML, GGUF specification, GPTQ, AWQ, QTIP and ExLlamaV3. Chapters 38-40 distinguish file interchange, repacked runtime layouts and native kernels.
- More models: Llama 3, YaRN, DeepSeek-V2 and DeepSeek-V3, together with their official model/config implementations.
- Images and retrieval: ViT, LLaVA, Qwen2-VL and Sentence-BERT. Chapter 43 implements a small vision path and pooling contracts; it does not validate trained VLM/retrieval quality.
- Many adapters: Punica and S-LoRA, for gathered low-rank kernels and adapter-aware memory management.
- Evaluation and reproducible launches: GSM8K, lm-evaluation-harness, vLLM serve, SGLang arguments, llama-server, TabbyAPI configuration and TensorRT-LLM serve. Chapter 44 provides launch recipes and a blank competitor-results template; a documentation check is not an executed benchmark.