F. Validation log
A book that teaches “test against a trusted reference” should say what it tested itself. This appendix records how the code and the numbers in this book were checked, in which environment, and what could not be checked there.
Policy
- Every program output shown in the book was produced by running the command shown, unless the text explicitly calls it illustrative or quotes another source.
- Every engine milestone is checked by tests that the reference implementation (
izh/) passes and that the generated skeleton (engine/) fails with aTODO(Chapter N)message. - Architectures are checked against independent reference implementations (Hugging Face Transformers) with random weights in FP32 before any real checkpoint is involved.
Completion environment (2026-10-05)
The original book validation ran on an x86-64 CPU-only environment. This completion pass ran on a DGX Spark with an available CUDA GPU; the results below separate these records.
| component | version / coverage |
|---|---|
| OS / CPU | Ubuntu 24.04.5 LTS, Linux aarch64 |
| GPU / driver | NVIDIA GB10, driver 580.178.04 |
| Python | 3.14.7 |
| PyTorch | 2.14.1+cu132, CUDA runtime 13.2 |
| CUDA compiler | nvcc 13.0.88 (/usr/local/cuda) |
| Triton | 3.8.0; native CUDA and a separate CPU-interpreter run |
| Transformers / tokenizers | 5.18.0 / 0.23.2 |
| FastAPI / httpx | 0.142.2 / 0.28.1 |
| GGUF reference package | 0.19.0 |
| Rust | rustc / cargo 1.99.0 |
| C++ / mdBook | g++ 13.3 / mdBook 0.5.4 |
No pretrained checkpoint was downloaded during this pass. The original environment reported blocked Hugging Face Hub access; current correctness checks use locally generated checkpoints, fixtures and random weights, rather than claiming real-checkpoint validation.
Completion results
| check | result |
|---|---|
IZH_IMPL=izh OMP_NUM_THREADS=2 pytest -rs (all chapters, CUDA available) | 374 passed, 2 skipped; skipped: flash_attn package unavailable and Metal requires an Apple GPU |
Chapters 38, 39, 43 and 44 with CUDA_VISIBLE_DEVICES='' TRITON_INTERPRET=1 IZH_DEVICE=cpu | 63 passed, 2 skipped; only the two explicit CUDA LoRA-kernel cases skip |
| generated skeleton spot checks | Chapter 43 M-RoPE and Chapter 44 arrival tests fail at their intended TODO(Chapter 43/44) functions |
tools/make_engine.py --check | all 87 generated package files match the reference/stub generator |
cargo test --release --offline | 18 unit + 2 checkpoint-fixture tests passed |
C++ Release build and izh all | 14 demos passed; CPU extension and native demo selected scalar kernels on this ARM build |
run.py features, run.py benchmark | ran on CPU; image/adapter/pooling plumbing, worked SLO arithmetic and independent greedy parity succeed |
| real loopback HTTP/SSE benchmark smoke | Poisson and burst runs, 8 requests each, 6 tokens/request, all 16 succeed; 48 output tokens per run; no competitor or production-performance claim |
| new HTTP contracts | mixed image + named LoRA matches offline generation; embedding float/base64, reranking, readiness, byte limits, deadline abort and slow-consumer cleanup pass |
| GGUF interoperability | legacy/K quant decoding, packed-to-affine repacking, reader/writer exchange and local model round trips checked with the upstream gguf package |
tools/check_book.py, mdbook build | 53 pages, 30 widget kinds; includes/links/anchors and HTML build pass |
prepare_book.py and lab sync | code download regenerated; source code artifacts mirrored to lab/inference-zero-to-hero-code, preserving local environment files |
The benchmark smoke uses a tiny random CPU model and a character-level test tokenizer. Its timings prove that requests, SSE parsing, usage, timestamps and shutdown work end to end. They are not representative throughput measurements and are not entered into Chapter 44’s competitor table. The quality smoke fixtures likewise do not establish trained retrieval, vision or GSM8K accuracy.
Earlier CPU-only records
The previous log recorded 144 passing and 13 GPU-skipped tests before Part VIII, 17 Rust unit tests plus 2 fixtures, and 13 C++ demos. It also recorded a full pre-Part-VIII skeleton run (125 expected failures, 5 fixture errors, 14 provided-code passes). Those historical counts are not the current suite size. The completion run checks generated-file equality and the two new chapter TODO boundaries rather than claiming another full skeleton run.
Parity with reference implementations (FP32, random weights)
| model / component | reference | max abs logit difference |
|---|---|---|
| GPT-2 | GPT2LMHeadModel | 1.8e-7 |
| Qwen3 (tied and untied heads, randomized norm weights) | Qwen3ForCausalLM | 3.0e-7 |
| Qwen3-MoE (with and without renormalization, 6 experts) | Qwen3MoeForCausalLM | 1.5e-7 |
| Gated DeltaNet layer | Qwen4ExpTextGatedDeltaNet | within the test’s 1e-5 tolerance |
| Qwen3.8-Flash-Next, 8 layers, all components, EOS mid-sequence | Qwen4ExpForCausalLM | 1.5e-8 (logits of magnitude 0.12) |
| Flash-Next with memory-mapped n-gram tables | own FP32 load | 0 (bit-identical) |
Bugs found by these checks
Recorded because each one is a lesson, and each is now covered by a test:
- Chunked Gated DeltaNet kept the wrong triangle of its system matrix. Chunk size 1 passed; sizes 5, 8 and 64 against the recurrent form caught it (Chapter 28).
- QSA top-k tie-breaking: masking invisible blocks with $-\infty$ in a longer vector broke ReLU-zero ties differently from the reference. Fixed by running
topkover exactly the visible blocks (Chapter 29). - Qwen3-MoE config: Transformers 5 writes
num_local_experts; the loader silently used its default of 8 experts and passed tests that happened to use 8. Fixed to accept both names and refuse missing fields; tests now use non-default sizes (Chapter 27). - safetensors:
torch.frombufferrejects zero-length tensors; empty tensors are now created directly (Chapter 9). - Full fine-tuning in BF16 with AdamW at learning rate 1e-5 loses most updates to rounding;
finetune.pynow keeps FP32 master weights with BF16 autocast for full fine-tuning (Chapters 13, 21). - Test collection: a test module that built models at import time turned one missing function into a collection error for the whole suite; models are now built inside tests (Chapter 16).
What was not validated here
Passing a teaching test establishes the tested contract and shapes, not every deployment or feature combination. The following boundaries remain explicit:
| path | completion status |
|---|---|
| CUDA kernels and Triton kernels | tiny correctness cases ran on GB10; broad architecture tuning and representative performance unmeasured |
| CUDA graphs and compile (Chapters 19, 33) | available GPU correctness tests passed; feature batches still take eager fallback, and production graph/memory tuning is unmeasured |
flash_attn adapter (Chapter 32) | skipped: package unavailable |
profile_num_blocks automatic GPU memory sizing | not independently exercised; tests specify pool sizes |
| GPU/CPU split and offload (Chapter 40) | tiny split-model correctness passed; representative transfer overlap and memory pressure unmeasured |
| NCCL, multiple physical GPUs, RDMA KV transfer | not run; Chapter 41 uses gloo CPU processes and local transfer contracts |
| Metal, ROCm, Vulkan, XPU | no target hardware validation; Metal test skipped; complete Vulkan/ROCm backends remain beyond this implementation |
| AVX2, AVX-512 VNNI, NEON dot-product kernels | not exercised by this ARM build; scalar CPU path passed. The original x86 record does not establish current cross-ISA coverage |
| real HF/GGUF checkpoints | no pretrained weights downloaded; architecture loaders tested with locally generated checkpoints and tiny fixtures |
| external GPTQ/AWQ/FP8/FP4 checkpoint tools, native EXL2/EXL3 | local synthetic layouts/reference equations tested; broad vendor/checkpoint interoperability not established; trellis lab is not native EXL3 loading |
| trained VLMs, embedding/retrieval quality, production rerankers | architecture and HTTP plumbing only; exact trained processors, towers, pair conventions and weights need independent validation |
| LoRA combinations | dense resident slots, mixed batches, cache keys and plain PEFT-layout loading tested; quantized/distributed/expert adapters, graph feature buffers and adapter-aware speculation are not validated |
| benchmark competitors | no vLLM, SGLang, llama.cpp, ExLlamaV3/TabbyAPI or TensorRT-LLM run; launch recipes checked against upstream docs, table intentionally blank |
| GSM8K/task accuracy | extractor and runner tested with fixtures; no pretrained-model benchmark score claimed |
| sustained load, device/process failures, orchestrated drain | unit/local HTTP checks only; multi-host recovery, rollout routing and long-duration SLO evidence require deployment tests |
| Chapter 1 probabilities, Chapter 18 chat transcript, expected GPU tables | retain their illustrative/source-attributed labels |
| fine-tuning/editing and bitsandbytes workflows | original tiny-checkpoint results retained; not rerun in this pass on real models |
| Flash-Next memory/deployment estimates | arithmetic, not measured serving of its pretrained weights |
The original CUDA compile-only and CPU-interpreter records remain historical evidence; current GPU correctness tests add coverage without turning estimates into measurements.
Reproducing
From the code/ directory:
IZH_IMPL=izh OMP_NUM_THREADS=2 pytest -rs # the reference, including available GPU paths
pytest # your engine: red until you implement it
(cd rust && cargo test --release)
cmake -S cpp -B build/cpp && cmake --build build/cpp -j && build/cpp/izh all
CUDA_VISIBLE_DEVICES='' TRITON_INTERPRET=1 IZH_DEVICE=cpu IZH_IMPL=izh pytest \
tests/test_ch38_gguf.py tests/test_ch39_quantized_checkpoints.py \
tests/test_ch43_multimodal_lora_embeddings.py tests/test_ch44_benchmarking.py
python run.py features --device cpu
python run.py benchmark --device cpu