D. Notation and glossary
Notation
| symbol | meaning |
|---|---|
| $B$ | batch size (sequences processed together) |
| $T$ | number of new tokens in a forward call (query length) |
| $S$ | number of keys visible (cached plus new) |
| $D$ or $d_\text{model}$ | hidden size (width of the residual stream) |
| $H$, $H_{kv}$ | query heads, key/value heads |
| $d$ or $D_h$ | head dimension |
| $V$ | vocabulary size |
| $L$ | number of layers |
| $E$, $k$ | number of experts, experts per token (Chapter 27); draft length (Chapter 26) |
| $r$, $\alpha$ | LoRA rank and scale (Chapter 22) |
| $W$ | a weight matrix, stored [out_features, in_features] as in nn.Linear |
| $\sigma$ | the logistic sigmoid $1/(1+e^{-x})$ |
| $\odot$ | elementwise product |
Tensor shapes are written in brackets, [B, T, D], with the fastest-varying axis last.
Glossary
Activation. Any intermediate value computed by a model during a forward pass; also the nonlinear functions (GELU, SiLU) applied in MLPs.
AdamW. The standard optimizer for transformers: per-parameter adaptive step sizes from running averages of the gradient and its square, with decoupled weight decay (Chapter 3).
Arithmetic intensity. FLOPs performed per byte moved from memory. Below the hardware’s ridge point an operation is memory-bound; above it, compute-bound (Chapter 10).
Attention. The operation that lets each position read from other positions: softmax-normalized query-key scores weight a sum of values (Chapter 5).
Autograd. Automatic differentiation by recording operations in a graph and applying the chain rule backwards (Chapter 3).
Bandwidth. Bytes per second a memory system delivers. Decode speed at small batch is bandwidth ÷ bytes per token (Chapters 1, 10).
Bank conflict. Several threads of a warp accessing different addresses in the same shared-memory bank, which serializes the accesses (Chapter 12).
BF16 (bfloat16). 16-bit floating point with FP32’s 8-bit exponent and a 7-bit mantissa: FP32’s range, about 3 significant digits (Chapter 13).
Block (CUDA). A group of threads that run on one SM, can share memory and synchronize with __syncthreads() (Chapter 11). In paged attention, a fixed-size chunk of the KV cache (Chapter 25).
BPE (byte-pair encoding). A tokenizer that starts from bytes and repeatedly merges the most frequent adjacent pair into a new token (Chapter 4).
Causal mask. The rule that position $t$ may attend only to positions $\le t$. This book implements it as key_pos <= query_pos, which also covers caches and chunks (Chapters 5, 16).
Chat template. The model-specific format that wraps conversation turns in special tokens (Chapter 4).
Chunked prefill. Splitting a long prompt into pieces processed over several steps, to bound per-step latency (Chapter 24).
Coalescing. Combining a warp’s memory accesses into few wide transactions when consecutive threads read consecutive addresses (Chapter 11).
Continuous batching. Scheduling at every decode step: finished requests leave and new ones join the running batch (Chapter 24).
Copy-on-write. Sharing a cache block between sequences until one writes to it, then copying (Chapter 25).
CUDA graph. A recorded sequence of GPU work replayed with one launch (Chapter 19).
Decode. Generating one token per step after prefill; memory-bound at small batch (Chapters 1, 16).
Delta rule. A state update that writes the error between a target value and what the memory currently returns for the key (Chapter 28).
Dequantization. Converting integer codes back to floating point with their scales (Chapter 20).
Embedding. A learned table mapping token IDs to vectors (Chapter 4).
Expert (MoE). One of several MLPs in a mixture-of-experts layer; a router chooses which experts process each token (Chapter 27).
FlashAttention. Exact attention computed tile by tile with an online softmax, never storing the full score matrix (Chapter 15).
FLOP. A floating-point operation; a multiply-add counts as 2.
Fusion. Combining several operations into one kernel so intermediates stay in registers or shared memory (Chapter 14).
Gated DeltaNet. A linear-attention layer with a decaying, delta-rule-updated matrix state (Chapter 28).
GQA (grouped-query attention). Several query heads share each key/value head, shrinking the KV cache (Chapters 5, 16).
Greedy decoding. Always choosing the highest-probability token (Chapter 8).
Hook. A function PyTorch runs after (or before) a module’s forward pass, able to read or replace its output (Chapter 23).
Hyper-connections. Several parallel residual streams with learned read and write weights per sublayer (Chapter 30).
KV cache. Stored keys and values of past positions, so each decode step computes only the new token (Chapter 16).
Kernel. A function that runs on the GPU, launched over a grid of threads or programs (Chapter 11).
Launch overhead. CPU time to issue a kernel; dominates when kernels are tiny (Chapter 19).
Layer normalization / RMSNorm. Rescaling each position’s vector to a fixed size; RMSNorm omits mean subtraction (Chapters 6, 17).
Logits. Unnormalized scores over the vocabulary; softmax turns them into probabilities (Chapter 1).
Logit lens. Applying the final norm and head to intermediate residual streams to see what each layer predicts (Chapter 23).
LoRA. Fine-tuning through a trainable low-rank update $\frac{\alpha}{r}BA$ added to frozen weights (Chapter 22).
Memory-bound / compute-bound. Limited by bytes moved, or by arithmetic throughput (Chapter 10).
MTP (multi-token prediction). Extra heads trained to predict tokens beyond the next; usable as speculative drafts (Chapters 26, 30).
N-gram memory. A hashed lookup table of embeddings for recent token pairs and triples, injected into the residual streams (Chapter 30).
NF4. A 4-bit data type with levels at quantiles of a normal distribution, used for QLoRA bases (Chapters 20, 22).
Occupancy. The fraction of an SM’s thread slots in use; higher occupancy hides memory latency (Chapter 12).
Online softmax. Computing softmax-weighted sums in one pass with a running maximum and rescaling (Chapter 15).
Paged attention. A KV cache in fixed-size blocks from a shared pool, addressed through per-sequence block tables (Chapter 25).
Perplexity. $e^{\text{loss}}$: the effective number of equally likely choices the model is uncertain between (Chapter 7).
Prefill. Processing the prompt in one forward pass to fill the cache and produce the first token’s logits; compute-bound (Chapters 1, 16).
Prefix caching. Reusing the KV cache of a shared prompt prefix across requests (Chapter 25).
QLoRA. LoRA on a frozen 4-bit (NF4) base model (Chapter 22).
QSA (Qwen Sparse Attention). Flash-Next’s block indexer that selects which keys full attention reads (Chapter 29).
Quantization. Storing values with fewer bits as integer codes plus scales (Chapter 20).
Residual stream. The running hidden state each layer reads from and adds to (Chapters 6, 23).
Roofline. The plot of attainable FLOP/s against arithmetic intensity: a bandwidth slope meeting a compute ceiling (Chapter 10).
RoPE (rotary position embedding). Encoding position by rotating query and key feature pairs by angles proportional to position (Chapter 17).
Safetensors. A checkpoint format with a JSON header and raw tensor bytes, safe to load (Chapter 9).
Sampling. Choosing the next token at random from a (possibly filtered) distribution: temperature, top-k, top-p, min-p (Chapter 8).
Shared memory. Fast on-chip memory shared by the threads of a block (Chapter 11).
SM (streaming multiprocessor). A GPU core: runs warps, holds registers and shared memory (Chapter 10).
Speculative decoding. A cheap draft proposes tokens and the target verifies them in one pass, with an exact acceptance rule (Chapter 26).
SwiGLU. The gated MLP used by Qwen and Llama: $W_\text{down}(\operatorname{SiLU}(xW_\text{gate}) \odot xW_\text{up})$ (Chapter 17).
Teacher forcing. Training on the true previous tokens rather than the model’s own outputs (Chapter 7).
Tensor core. Hardware that multiplies small matrix tiles per instruction, at much higher throughput than ordinary arithmetic (Chapter 12).
Tiling. Splitting a computation into blocks that fit in fast memory, to reuse each loaded value many times (Chapter 12).
Token. One unit of the tokenizer’s vocabulary; the model sees only token IDs (Chapter 4).
TTFT / TPOT. Time to first token; time per output token (Chapters 18, 24).
Warp. 32 threads that execute together on NVIDIA GPUs (Chapter 11).
Weight tying. Using the token embedding matrix as the output head (Chapters 6, 9).
Serving, formats and model coverage
Adapter slot / lease. Resident LoRA storage and a request’s right to keep its contents unchanged until completion; cache identity comes from effective weights, not the slot (Chapter 43).
AWQ / GPTQ. Activation-aware column scaling and curvature-aware weight-error compensation, respectively, with checkpoint-specific packed layouts (Chapter 39).
BGMV / SGMV. Batched or segmented gathered matrix-vector multiplication: low-rank shrink/expand operations selected by adapter per row or row segment (Chapter 43).
Chunked prefill. Computing a prompt in bounded token chunks mixed with other requests’ decode work (Chapters 24, 31).
Disaggregated prefill/decode. Separate engines for prompt processing and generation, with KV and request state transferred between them (Chapter 41).
E2E / ITL. End-to-end request latency and inter-token latency. Streamed chunk gaps are reported separately when token arrival times are unknown (Chapter 44).
Expert / tensor / pipeline / context parallelism. Splitting experts, projection matrices, layer stages or context positions across devices, with different communication patterns (Chapter 41).
FP8 block / MXFP4 / NVFP4. Floating-point quantization with block scales; the formats differ in value bits, scale encoding and block size (Chapter 39).
GGUF / GGML quants. A model file containing metadata and tensors, and the tensor block formats commonly carried inside it (Chapter 38).
Goodput / SLO. Successful requests per second meeting all configured service-level objectives; failures remain in attainment’s denominator (Chapter 44).
Guided decoding. Masking tokens according to a byte-level constraint automaton, enforcing a supported regex or schema language (Chapter 34).
Hash-chained prefix key. A fixed-size cryptographic digest of preceding blocks, current tokens and activation-changing inputs such as images and adapters (Chapters 31, 43).
MLA. Multi-head latent attention: cached compressed latent vectors plus shared rotary keys, with head-specific projections absorbed into queries and outputs during decode (Chapter 42).
M-RoPE. Rotary frequency pairs assigned to time, height and width coordinates; text uses equal coordinates on every axis (Chapter 43).
Open-loop arrivals. Request submission times chosen independently of earlier completions; used to expose queueing and overload (Chapter 44).
Pooling / cross-encoder reranking. Producing one vector from token hidden states, or scoring jointly encoded query-document pairs with a trained head (Chapter 43).
Preemption. Releasing a request’s device KV by dropping it for recomputation or saving it in host memory for later restoration (Chapter 31).
Ragged batch. One flattened token matrix whose sequence boundaries, lengths and cache slots are described by metadata (Chapter 31).
Readiness / drain. Whether a process can accept new work, and the shutdown phase that stops admission while existing work completes under a deadline (Chapter 44).
Split-KV. Partitioning long-context attention across programs, then merging outputs using each partition’s log-sum-exp normalizer (Chapter 32).
Trellis quantization. A finite-state sequence of codes optimized jointly over a weight vector, rather than independent scalar rounding (Chapter 39).