Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

F. Validation log

A book that teaches “test against a trusted reference” should say what it tested itself. This appendix records how the code and the numbers in this book were checked, in which environment, and what could not be checked there.

Policy

  • Every program output shown in the book was produced by running the command shown, unless the text explicitly calls it illustrative or quotes another source.
  • Every engine milestone is checked by tests that the reference implementation (izh/) passes and that the generated skeleton (engine/) fails with a TODO(Chapter N) message.
  • Architectures are checked against independent reference implementations (Hugging Face Transformers) with random weights in FP32 before any real checkpoint is involved.

Completion environment (2026-10-05)

The original book validation ran on an x86-64 CPU-only environment. This completion pass ran on a DGX Spark with an available CUDA GPU; the results below separate these records.

componentversion / coverage
OS / CPUUbuntu 24.04.5 LTS, Linux aarch64
GPU / driverNVIDIA GB10, driver 580.178.04
Python3.14.7
PyTorch2.14.1+cu132, CUDA runtime 13.2
CUDA compilernvcc 13.0.88 (/usr/local/cuda)
Triton3.8.0; native CUDA and a separate CPU-interpreter run
Transformers / tokenizers5.18.0 / 0.23.2
FastAPI / httpx0.142.2 / 0.28.1
GGUF reference package0.19.0
Rustrustc / cargo 1.99.0
C++ / mdBookg++ 13.3 / mdBook 0.5.4

No pretrained checkpoint was downloaded during this pass. The original environment reported blocked Hugging Face Hub access; current correctness checks use locally generated checkpoints, fixtures and random weights, rather than claiming real-checkpoint validation.

Completion results

checkresult
IZH_IMPL=izh OMP_NUM_THREADS=2 pytest -rs (all chapters, CUDA available)374 passed, 2 skipped; skipped: flash_attn package unavailable and Metal requires an Apple GPU
Chapters 38, 39, 43 and 44 with CUDA_VISIBLE_DEVICES='' TRITON_INTERPRET=1 IZH_DEVICE=cpu63 passed, 2 skipped; only the two explicit CUDA LoRA-kernel cases skip
generated skeleton spot checksChapter 43 M-RoPE and Chapter 44 arrival tests fail at their intended TODO(Chapter 43/44) functions
tools/make_engine.py --checkall 87 generated package files match the reference/stub generator
cargo test --release --offline18 unit + 2 checkpoint-fixture tests passed
C++ Release build and izh all14 demos passed; CPU extension and native demo selected scalar kernels on this ARM build
run.py features, run.py benchmarkran on CPU; image/adapter/pooling plumbing, worked SLO arithmetic and independent greedy parity succeed
real loopback HTTP/SSE benchmark smokePoisson and burst runs, 8 requests each, 6 tokens/request, all 16 succeed; 48 output tokens per run; no competitor or production-performance claim
new HTTP contractsmixed image + named LoRA matches offline generation; embedding float/base64, reranking, readiness, byte limits, deadline abort and slow-consumer cleanup pass
GGUF interoperabilitylegacy/K quant decoding, packed-to-affine repacking, reader/writer exchange and local model round trips checked with the upstream gguf package
tools/check_book.py, mdbook build53 pages, 30 widget kinds; includes/links/anchors and HTML build pass
prepare_book.py and lab synccode download regenerated; source code artifacts mirrored to lab/inference-zero-to-hero-code, preserving local environment files

The benchmark smoke uses a tiny random CPU model and a character-level test tokenizer. Its timings prove that requests, SSE parsing, usage, timestamps and shutdown work end to end. They are not representative throughput measurements and are not entered into Chapter 44’s competitor table. The quality smoke fixtures likewise do not establish trained retrieval, vision or GSM8K accuracy.

Earlier CPU-only records

The previous log recorded 144 passing and 13 GPU-skipped tests before Part VIII, 17 Rust unit tests plus 2 fixtures, and 13 C++ demos. It also recorded a full pre-Part-VIII skeleton run (125 expected failures, 5 fixture errors, 14 provided-code passes). Those historical counts are not the current suite size. The completion run checks generated-file equality and the two new chapter TODO boundaries rather than claiming another full skeleton run.

Parity with reference implementations (FP32, random weights)

model / componentreferencemax abs logit difference
GPT-2GPT2LMHeadModel1.8e-7
Qwen3 (tied and untied heads, randomized norm weights)Qwen3ForCausalLM3.0e-7
Qwen3-MoE (with and without renormalization, 6 experts)Qwen3MoeForCausalLM1.5e-7
Gated DeltaNet layerQwen4ExpTextGatedDeltaNetwithin the test’s 1e-5 tolerance
Qwen3.8-Flash-Next, 8 layers, all components, EOS mid-sequenceQwen4ExpForCausalLM1.5e-8 (logits of magnitude 0.12)
Flash-Next with memory-mapped n-gram tablesown FP32 load0 (bit-identical)

Bugs found by these checks

Recorded because each one is a lesson, and each is now covered by a test:

  1. Chunked Gated DeltaNet kept the wrong triangle of its system matrix. Chunk size 1 passed; sizes 5, 8 and 64 against the recurrent form caught it (Chapter 28).
  2. QSA top-k tie-breaking: masking invisible blocks with $-\infty$ in a longer vector broke ReLU-zero ties differently from the reference. Fixed by running topk over exactly the visible blocks (Chapter 29).
  3. Qwen3-MoE config: Transformers 5 writes num_local_experts; the loader silently used its default of 8 experts and passed tests that happened to use 8. Fixed to accept both names and refuse missing fields; tests now use non-default sizes (Chapter 27).
  4. safetensors: torch.frombuffer rejects zero-length tensors; empty tensors are now created directly (Chapter 9).
  5. Full fine-tuning in BF16 with AdamW at learning rate 1e-5 loses most updates to rounding; finetune.py now keeps FP32 master weights with BF16 autocast for full fine-tuning (Chapters 13, 21).
  6. Test collection: a test module that built models at import time turned one missing function into a collection error for the whole suite; models are now built inside tests (Chapter 16).

What was not validated here

Passing a teaching test establishes the tested contract and shapes, not every deployment or feature combination. The following boundaries remain explicit:

pathcompletion status
CUDA kernels and Triton kernelstiny correctness cases ran on GB10; broad architecture tuning and representative performance unmeasured
CUDA graphs and compile (Chapters 19, 33)available GPU correctness tests passed; feature batches still take eager fallback, and production graph/memory tuning is unmeasured
flash_attn adapter (Chapter 32)skipped: package unavailable
profile_num_blocks automatic GPU memory sizingnot independently exercised; tests specify pool sizes
GPU/CPU split and offload (Chapter 40)tiny split-model correctness passed; representative transfer overlap and memory pressure unmeasured
NCCL, multiple physical GPUs, RDMA KV transfernot run; Chapter 41 uses gloo CPU processes and local transfer contracts
Metal, ROCm, Vulkan, XPUno target hardware validation; Metal test skipped; complete Vulkan/ROCm backends remain beyond this implementation
AVX2, AVX-512 VNNI, NEON dot-product kernelsnot exercised by this ARM build; scalar CPU path passed. The original x86 record does not establish current cross-ISA coverage
real HF/GGUF checkpointsno pretrained weights downloaded; architecture loaders tested with locally generated checkpoints and tiny fixtures
external GPTQ/AWQ/FP8/FP4 checkpoint tools, native EXL2/EXL3local synthetic layouts/reference equations tested; broad vendor/checkpoint interoperability not established; trellis lab is not native EXL3 loading
trained VLMs, embedding/retrieval quality, production rerankersarchitecture and HTTP plumbing only; exact trained processors, towers, pair conventions and weights need independent validation
LoRA combinationsdense resident slots, mixed batches, cache keys and plain PEFT-layout loading tested; quantized/distributed/expert adapters, graph feature buffers and adapter-aware speculation are not validated
benchmark competitorsno vLLM, SGLang, llama.cpp, ExLlamaV3/TabbyAPI or TensorRT-LLM run; launch recipes checked against upstream docs, table intentionally blank
GSM8K/task accuracyextractor and runner tested with fixtures; no pretrained-model benchmark score claimed
sustained load, device/process failures, orchestrated drainunit/local HTTP checks only; multi-host recovery, rollout routing and long-duration SLO evidence require deployment tests
Chapter 1 probabilities, Chapter 18 chat transcript, expected GPU tablesretain their illustrative/source-attributed labels
fine-tuning/editing and bitsandbytes workflowsoriginal tiny-checkpoint results retained; not rerun in this pass on real models
Flash-Next memory/deployment estimatesarithmetic, not measured serving of its pretrained weights

The original CUDA compile-only and CPU-interpreter records remain historical evidence; current GPU correctness tests add coverage without turning estimates into measurements.

Reproducing

From the code/ directory:

IZH_IMPL=izh OMP_NUM_THREADS=2 pytest -rs  # the reference, including available GPU paths
pytest                                     # your engine: red until you implement it
(cd rust && cargo test --release)
cmake -S cpp -B build/cpp && cmake --build build/cpp -j && build/cpp/izh all
CUDA_VISIBLE_DEVICES='' TRITON_INTERPRET=1 IZH_DEVICE=cpu IZH_IMPL=izh pytest \
  tests/test_ch38_gguf.py tests/test_ch39_quantized_checkpoints.py \
  tests/test_ch43_multimodal_lora_embeddings.py tests/test_ch44_benchmarking.py
python run.py features --device cpu
python run.py benchmark --device cpu