Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

45. Where to go next

In this chapter

  • Identify what the completed engine implements and where its boundaries remain.
  • Use a repeatable bring-up method for new models, hardware and feature combinations.
  • Choose a concrete next project using correctness evidence and measured bottlenecks.

Time: as long as you like.

What you’ve built

You began with tensors, an autograd engine and a tokenizer. You trained a GPT, loaded real model formats, wrote CUDA and Triton kernels, built cached decoding, changed models with fine-tuning and LoRA, and matched a frontier hybrid architecture to its reference. Part VIII then put serving mechanisms together: flattened mixed batches, paging and prefix caching, preemption, graph buckets, batched sampling, text boundaries, an asynchronous HTTP frontend, batched speculation, quantized checkpoint formats, offloading, distributed execution, model registration, latent attention, image inputs, pooling and many adapters. Chapter 44 added the client that measures the resulting service.

The most valuable result is the method: derive, implement the simple version, compare with an independent reference, optimize while keeping the comparison green, then measure the system. That method transfers to a model or device this book has never seen.

There is still a difference between implementing a mechanism and operating a dependable inference service. Tests on tiny random models establish useful invariants. They do not establish real-checkpoint coverage, tail latency under sustained load or recovery after a GPU fails. Appendix F records the evidence this book actually has.

What remains beyond the book

The earlier final chapter left distributed serving as a next project. Chapter 41 now implements its mechanics; Chapter 42 brings up more model families; Chapter 43 adds non-text boundaries. The next work is deeper coverage and integration:

areabuilt herenext level
attention and matmulpaged unified/split-KV kernels, affine quantized matmul, FP8 weight readsarchitecture-specific TMA/WGMMA pipelines, optimized MLA prefill/decode, tensor-core INT4/FP8/FP4 kernels and autotuning
formatsGGUF/selected GGML quants, GPTQ/AWQ layouts, block FP8, MX/NV FP4 references, a trellis labbroad versioned interoperability, every GGML/i-quant variant and native EXL2/EXL3 loading; checkpoint-specific quality validation
hardwareCPU references and selected SIMD paths, CUDA/Triton, a Metal lab and platform detectioncomplete optimized Metal/Vulkan/ROCm backends, kernel tuning per device, reliable heterogeneous deployment
many GPUsTP/PP/EP, ring attention, routing and KV transfer reference pathsoverlapping communication, topology-aware collectives, RDMA KV transport, resharding and distributed failure recovery
modelsdense registry, MoE and MLA, a separate hybrid capstone, a tiny vision pathtrained multimodal checkpoint loaders, audio/video processors, native batched hybrid-state scheduling and broader model contracts
LoRA and poolingnamed resident slots, gathered kernels, bounded embedding/reranking forwardsadapter paging, quantized/distributed adapters, graph feature buffers and cross-request embedding batching
servicecancellation, limits, deadlines, readiness, drain, metrics and load generatortenant identity/quotas, rollout routing, durable telemetry, availability targets and fault-tested operations
combinationsshared paged scheduler with tested feature pathsa measured compatibility matrix for graphs, speculation, guides, quantization, LoRA, images, offload and parallelism together

That last row matters. An adapter request currently takes an eager path, and a feature request with a drafter is refused. The capstone’s recurrent state is not automatically paged like dense K/V. A reference matmul reading quantized weights is not a fused low-bit tensor-core implementation. Read the explicit limitations in each chapter; don’t advertise a combination just because each feature exists in isolation.

Bring up a new architecture

Use the checklist that brought up the book’s model families:

  1. Read the config and reference code. Build Chapter 30’s ledger: equations, names, shapes, dtypes, masks, positions, state and trained input/output conventions.
  2. List behavior changes. A different vocabulary size is a parameter; a different norm, rotary layout, router or image processor is a new computation.
  3. Reject unsupported options. Never let a config field that changes behavior quietly fall through to a default. Add a test for the refusal too.
  4. Test components independently. Randomize all weights, including norms and gates that initialize to zero or one; use non-default dimensions and odd lengths.
  5. Save and reload a tiny reference checkpoint. Compare FP32 hidden states/logits at each layer, then at deployment precision. Check every tensor was consumed or explicitly excluded.
  6. Test state equivalence. Full forward, chunked prefill, decode, prefix reuse and preemption should describe the same history. For hybrid states, test snapshots and restoration.
  7. Test feature combinations. Mix request lengths and adapters; interrupt an image chunk; force memory pressure; cancel a request with shared prefixes; change batch membership. Pick invariants that catch errors rather than repeating implementation steps.
  8. Plan memory and load real weights. Account for activations, graph buffers, scale tensors, adapter slots and host copies as well as weights and KV. Compare real greedy outputs with a trusted implementation.
  9. Measure the service. Use Chapter 44’s workloads, a declared accuracy protocol and the hardware ceiling. Optimize the bottleneck you measured, then repeat the correctness check and experiment.

Read production engines as a request journey

Trace one request from HTTP parsing through scheduling, cache allocation, the model’s forward, attention dispatch, sampling and output. You now have names for each boundary:

engineuseful entry pointbook counterpart
vLLMV1 scheduler, KV manager, model runner and attention backendsChapters 31-34 and 42
SGLangtokenizer manager, scheduler and radix cacheChapters 31, 35-37
llama.cppmodel loader, GGML graph and backend kernelsChapters 38-40
ExLlamaV3loader, linear modules, cache and generatorChapters 25, 37, 39 and 44
TensorRT-LLMLLM API, executor and kernel dispatchChapters 32-33, 39 and 41

Start from a pinned revision, since module paths move. Chapter 44 links the projects’ current launch documentation. Find where the real engine handles a corner case your reference refuses: adapter rank changes, graph fallback, a cache eviction during speculation, or a failed KV transfer. That is a good first contribution because you can state the invariant and write a focused regression test.

Open questions worth experimenting with

  • Long context with less state. Compare sparse retrieval, recurrent/hybrid state and selective KV retention on recall tasks, not just cache size. Which information is lost, and at what context length?
  • Low-bit everything. Study weight, activation and KV quantization separately before combining them. Add quality, scale-storage bytes and kernel traffic to the performance curve.
  • Speculation for complex state. Tree verification, MTP and hybrid-state rollback need ownership and commit rules as much as faster draft models. Can an adapter-aware drafter stay useful across many adapters?
  • Batch-invariant inference. Results should be reproducible under a declared contract when neighbouring requests change. Measure the throughput cost of deterministic reductions and routing.
  • Scheduling heterogeneous work. An image encoder, an embedding batch, long prefill and a decode step have different resource footprints. Design admission and scheduling that protect decode SLOs without starving the other work.
  • Kernel specialization with a fallback. Optimize one costly architecture/device pair while keeping a simple reference that verifies every supported shape and supplies correct results elsewhere.

Follow GPU Mode, model technical reports and the systems projects you use. Read release code with a specific question, then reproduce one result. A small measured experiment is more useful than collecting every new acronym.

A last exercise

Pick one concrete boundary in the table above. Bring up a trained VLM, capture adapter batches in graphs, move KV over an actual RDMA link, or optimize a quantized matmul on your device. Write the contract first. Match a trusted reference, exercise failure and state transitions, save a reproducible benchmark with accuracy beside speed, and publish the result with its limitations.

Where: create experiments/final_project/ in the code root for the contract, scripts and results, with a README.md for the write-up. For the implementation, start with the targets in Chapter 43 for a VLM or adapter graphs, Chapter 41 for KV transfer, or Chapter 39 for a quantized matmul.

That write-up is the proof that you can continue without a book telling you the next step.