45. Where to go next
In this chapter
- Identify what the completed engine implements and where its boundaries remain.
- Use a repeatable bring-up method for new models, hardware and feature combinations.
- Choose a concrete next project using correctness evidence and measured bottlenecks.
Time: as long as you like.
What you’ve built
You began with tensors, an autograd engine and a tokenizer. You trained a GPT, loaded real model formats, wrote CUDA and Triton kernels, built cached decoding, changed models with fine-tuning and LoRA, and matched a frontier hybrid architecture to its reference. Part VIII then put serving mechanisms together: flattened mixed batches, paging and prefix caching, preemption, graph buckets, batched sampling, text boundaries, an asynchronous HTTP frontend, batched speculation, quantized checkpoint formats, offloading, distributed execution, model registration, latent attention, image inputs, pooling and many adapters. Chapter 44 added the client that measures the resulting service.
The most valuable result is the method: derive, implement the simple version, compare with an independent reference, optimize while keeping the comparison green, then measure the system. That method transfers to a model or device this book has never seen.
There is still a difference between implementing a mechanism and operating a dependable inference service. Tests on tiny random models establish useful invariants. They do not establish real-checkpoint coverage, tail latency under sustained load or recovery after a GPU fails. Appendix F records the evidence this book actually has.
What remains beyond the book
The earlier final chapter left distributed serving as a next project. Chapter 41 now implements its mechanics; Chapter 42 brings up more model families; Chapter 43 adds non-text boundaries. The next work is deeper coverage and integration:
| area | built here | next level |
|---|---|---|
| attention and matmul | paged unified/split-KV kernels, affine quantized matmul, FP8 weight reads | architecture-specific TMA/WGMMA pipelines, optimized MLA prefill/decode, tensor-core INT4/FP8/FP4 kernels and autotuning |
| formats | GGUF/selected GGML quants, GPTQ/AWQ layouts, block FP8, MX/NV FP4 references, a trellis lab | broad versioned interoperability, every GGML/i-quant variant and native EXL2/EXL3 loading; checkpoint-specific quality validation |
| hardware | CPU references and selected SIMD paths, CUDA/Triton, a Metal lab and platform detection | complete optimized Metal/Vulkan/ROCm backends, kernel tuning per device, reliable heterogeneous deployment |
| many GPUs | TP/PP/EP, ring attention, routing and KV transfer reference paths | overlapping communication, topology-aware collectives, RDMA KV transport, resharding and distributed failure recovery |
| models | dense registry, MoE and MLA, a separate hybrid capstone, a tiny vision path | trained multimodal checkpoint loaders, audio/video processors, native batched hybrid-state scheduling and broader model contracts |
| LoRA and pooling | named resident slots, gathered kernels, bounded embedding/reranking forwards | adapter paging, quantized/distributed adapters, graph feature buffers and cross-request embedding batching |
| service | cancellation, limits, deadlines, readiness, drain, metrics and load generator | tenant identity/quotas, rollout routing, durable telemetry, availability targets and fault-tested operations |
| combinations | shared paged scheduler with tested feature paths | a measured compatibility matrix for graphs, speculation, guides, quantization, LoRA, images, offload and parallelism together |
That last row matters. An adapter request currently takes an eager path, and a feature request with a drafter is refused. The capstone’s recurrent state is not automatically paged like dense K/V. A reference matmul reading quantized weights is not a fused low-bit tensor-core implementation. Read the explicit limitations in each chapter; don’t advertise a combination just because each feature exists in isolation.
Bring up a new architecture
Use the checklist that brought up the book’s model families:
- Read the config and reference code. Build Chapter 30’s ledger: equations, names, shapes, dtypes, masks, positions, state and trained input/output conventions.
- List behavior changes. A different vocabulary size is a parameter; a different norm, rotary layout, router or image processor is a new computation.
- Reject unsupported options. Never let a config field that changes behavior quietly fall through to a default. Add a test for the refusal too.
- Test components independently. Randomize all weights, including norms and gates that initialize to zero or one; use non-default dimensions and odd lengths.
- Save and reload a tiny reference checkpoint. Compare FP32 hidden states/logits at each layer, then at deployment precision. Check every tensor was consumed or explicitly excluded.
- Test state equivalence. Full forward, chunked prefill, decode, prefix reuse and preemption should describe the same history. For hybrid states, test snapshots and restoration.
- Test feature combinations. Mix request lengths and adapters; interrupt an image chunk; force memory pressure; cancel a request with shared prefixes; change batch membership. Pick invariants that catch errors rather than repeating implementation steps.
- Plan memory and load real weights. Account for activations, graph buffers, scale tensors, adapter slots and host copies as well as weights and KV. Compare real greedy outputs with a trusted implementation.
- Measure the service. Use Chapter 44’s workloads, a declared accuracy protocol and the hardware ceiling. Optimize the bottleneck you measured, then repeat the correctness check and experiment.
Read production engines as a request journey
Trace one request from HTTP parsing through scheduling, cache allocation, the model’s forward, attention dispatch, sampling and output. You now have names for each boundary:
| engine | useful entry point | book counterpart |
|---|---|---|
| vLLM | V1 scheduler, KV manager, model runner and attention backends | Chapters 31-34 and 42 |
| SGLang | tokenizer manager, scheduler and radix cache | Chapters 31, 35-37 |
| llama.cpp | model loader, GGML graph and backend kernels | Chapters 38-40 |
| ExLlamaV3 | loader, linear modules, cache and generator | Chapters 25, 37, 39 and 44 |
| TensorRT-LLM | LLM API, executor and kernel dispatch | Chapters 32-33, 39 and 41 |
Start from a pinned revision, since module paths move. Chapter 44 links the projects’ current launch documentation. Find where the real engine handles a corner case your reference refuses: adapter rank changes, graph fallback, a cache eviction during speculation, or a failed KV transfer. That is a good first contribution because you can state the invariant and write a focused regression test.
Open questions worth experimenting with
- Long context with less state. Compare sparse retrieval, recurrent/hybrid state and selective KV retention on recall tasks, not just cache size. Which information is lost, and at what context length?
- Low-bit everything. Study weight, activation and KV quantization separately before combining them. Add quality, scale-storage bytes and kernel traffic to the performance curve.
- Speculation for complex state. Tree verification, MTP and hybrid-state rollback need ownership and commit rules as much as faster draft models. Can an adapter-aware drafter stay useful across many adapters?
- Batch-invariant inference. Results should be reproducible under a declared contract when neighbouring requests change. Measure the throughput cost of deterministic reductions and routing.
- Scheduling heterogeneous work. An image encoder, an embedding batch, long prefill and a decode step have different resource footprints. Design admission and scheduling that protect decode SLOs without starving the other work.
- Kernel specialization with a fallback. Optimize one costly architecture/device pair while keeping a simple reference that verifies every supported shape and supplies correct results elsewhere.
Follow GPU Mode, model technical reports and the systems projects you use. Read release code with a specific question, then reproduce one result. A small measured experiment is more useful than collecting every new acronym.
A last exercise
Pick one concrete boundary in the table above. Bring up a trained VLM, capture adapter batches in graphs, move KV over an actual RDMA link, or optimize a quantized matmul on your device. Write the contract first. Match a trusted reference, exercise failure and state transitions, save a reproducible benchmark with accuracy beside speed, and publish the result with its limitations.
Where: create experiments/final_project/ in the code root for the contract, scripts and
results, with a README.md for the write-up. For the implementation, start with the targets in
Chapter 43 for a VLM or adapter graphs,
Chapter 41 for KV transfer, or
Chapter 39 for a quantized matmul.
That write-up is the proof that you can continue without a book telling you the next step.