Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

44. Benchmark the whole engine, then harden it

In this chapter

  • Measure user latency and throughput under an arrival process, rather than timing an isolated forward pass.
  • Separate TTFT, TPOT, token ITL, streamed chunk gaps and end-to-end latency.
  • Count goodput under explicit service-level objectives, with failures kept visible.
  • Compare greedy trajectories, perplexity, KL and a small task evaluation before interpreting speed.
  • Launch the same checkpoint in the book's engine and production engines, with a reproducible workload and an honest results table.
  • Bound request bodies, deadlines, output buffers and shutdown; distinguish health from readiness.

You will build

arrival_offsets, request and summarize in engine/bench.py; logit_quality and evaluate_endpoint in engine/evaluation.py.

Time: 5-7 hours plus experiments. GPU: optional for the milestone; meaningful engine comparisons need representative hardware and real weights.

What are we trying to make faster?

Chapter 10 timed kernels. Chapter 19 timed one decoder. A user sends an HTTP request, waits in admission and scheduling queues, waits for prefill, then receives tokens. A kernel improvement matters only if it improves that experience, or lets the service handle more users at the same experience.

Consider two engines. One produces 1,000 tokens/s but stalls some requests for several seconds. Another produces 800 tokens/s and keeps every request below the latency target. Which is better depends on the service’s objective. Report throughput and latency together, at several offered loads, and identify the load at which the target starts failing.

Before timing, write an experiment record: hardware and power limits, engine revision, model/tokenizer revision, precision, quantization, cache dtype and capacity, context limit, parallelism, graph and speculation settings, prefix reuse policy, prompt/output length distributions, sampling settings and arrival process. --metadata attaches that JSON record to every result. Record the full launch command too. A model name and a tokens/s number are not a reproducible benchmark.

One request’s clocks

Use a monotonic clock on the client. Let $a$ be scheduled arrival, $d$ HTTP dispatch, $f$ first visible content, $l$ last visible content, $e$ completed stream, and $n$ the generated token count:

metricdefinitionwhat it captures
client queue$d-a$waiting for the load generator’s connection/concurrency cap
TTFT$f-a$client queue, transport, engine queue, prefill and first visible output
E2E$e-a$everything through stream termination
TPOT$(l-f)/(n-1)$, if $n>1$average time per output token after the first
ITLgap between consecutive token arrivalsirregular decode progress
chunk gapgap between consecutive content eventswhat the streaming client actually observed

An SSE role header, keepalive or empty finish event is not the first token. The benchmark waits for nonempty content. A Unicode character may need several token bytes before becoming visible, and a speculative step may stream several tokens together. TTFT here is first visible text, which can differ from an engine’s first-token-ID timestamp. Keep that distinction when comparing server metrics with client metrics.

A chunk is not necessarily a token. The client records content-event timestamps and token counts from final usage. It reports token ITL only when logprobs prove that each event contains exactly one token and the total equals usage. Otherwise the ITL sample is empty, while chunk-gap percentiles remain available. It cannot recover true individual token arrival times by dividing a five-token chunk into five imaginary arrivals.

TPOT uses total token count and the first/last content times. If output arrives entirely in one chunk, it is unknown rather than zero. Usage is preferred; an optional tokenizer fallback retokenizes the visible text and labels the count as an estimate. Stop tokens, special tokens and invisible reasoning can make that estimate differ from the engine’s generated-token count.

An open-loop load generator

In a closed loop, a client submits another request when one finishes. A slower engine then sees less offered load, which hides its overload behavior. In an open loop, arrivals are chosen independently of completions:

  • Constant arrivals are spaced by $1/\lambda$ seconds.
  • Poisson arrivals use exponential gaps with mean $1/\lambda$; bursts arise naturally.
  • Explicit bursts send $b$ requests together every $b/\lambda$ seconds.
def arrival_offsets(count, rate, kind="poisson", seed=0, burst_size=8):
    """Open loop: arrivals don't wait for completions. Poisson gaps have mean 1/rate.
    Burst groups arrive together, with the same long-run offered rate.  (Your engine: Chapter 44)
    """
    if count < 1 or rate <= 0 or math.isnan(rate) or burst_size < 1 or kind not in ("poisson", "burst", "constant"):
        raise ValueError("Need positive count/rate/burst_size and a known arrival process")
    if math.isinf(rate):
        return [0.0] * count
    rng, at, out = random.Random(seed), 0.0, [0.0]
    for i in range(1, count):
        if kind == "poisson":
            at += rng.expovariate(rate)
        elif kind == "constant":
            at = i / rate
        elif i % burst_size == 0:
            at += burst_size / rate
        out.append(at)
    return out

The first request arrives at zero; rate=inf sends the finite workload together. A seeded random generator makes arrival times reproducible. A semaphore caps simultaneous HTTP connections, but scheduled arrival stays unchanged, so local queueing is included in TTFT and E2E. Inspect client_queue: if it grows, the client cap is part of your result. Raise it, or declare that you are measuring a client-concurrency-limited service.

async def request(client, url, model, workload, trace, clock=time.perf_counter, tokenizer=None, api_key=None):
    """Time actual content events, require [DONE], retain final usage, and record errors.
    The logprobs length can prove token/event counts; otherwise ITL is chunk-level.  (Your engine: Chapter 44)
    """
    payload = {"model": model, "prompt": workload["prompt"], "max_tokens": workload.get("max_tokens", 128),
               "temperature": 0, "n": 1, "stream": True, "stream_options": {"include_usage": True}}
    headers = {"Authorization": f"Bearer {api_key}"} if api_key else {}
    done = False
    try:
        async with client.stream("POST", url, json=payload, headers=headers) as response:
            response.raise_for_status()
            async for data in sse_events(response.aiter_lines()):
                if data == "[DONE]":
                    done = True
                    break
                event = json.loads(data)
                if event.get("error"):
                    raise ValueError(f"SSE error: {event['error']}")
                usage = event.get("usage")
                if usage:
                    trace.prompt_tokens, trace.output_tokens = usage.get("prompt_tokens"), usage.get("completion_tokens")
                    trace.token_count_source = "usage"
                for choice in event.get("choices", []):
                    if choice.get("index", 0) != 0:
                        continue
                    text = choice.get("text", "")
                    delta = choice.get("delta") or {}
                    text += delta.get("content") or delta.get("reasoning_content") or ""
                    if text:
                        trace.content_times.append(clock())
                        trace.text += text
                        lp = choice.get("logprobs") or {}
                        tokens = lp.get("tokens", lp.get("content"))
                        trace.event_token_counts.append(len(tokens) if tokens is not None else None)
        if not done:
            raise ValueError("Stream ended without [DONE]")
        if trace.output_tokens is None and tokenizer is not None:
            trace.output_tokens = len(tokenizer.encode(trace.text))
            trace.token_count_source = "retokenized_text"
        if trace.output_tokens is not None and (type(trace.output_tokens) is not int or trace.output_tokens < 0):
            raise ValueError("Invalid completion_tokens")
        if not trace.content_times:
            raise ValueError("No generated content received")
    except asyncio.CancelledError:
        raise
    except Exception as error:
        trace.error = f"{type(error).__name__}: {error}"
    finally:
        trace.finished = clock()
    return trace

httpx.aiter_lines handles fragmented network bytes and UTF-8 decoding. sse_events assembles multiline data fields, ignores comments and recognizes [DONE]. A transport success without [DONE], a server error event, an HTTP 503 or a timeout becomes an explicit failed trace. Warmup failures stop the experiment. Requests have both transport timeouts and a total deadline; a stream that periodically sends bytes cannot run forever.

The JSON output preserves workload, configuration and every request’s timestamps, token counts, text and error. Those traces let you recompute percentiles or change SLOs after a run without sending the workload again. They may contain sensitive prompts and outputs; use a synthetic workload when sharing them.

Throughput and goodput

The measurement window begins at the first scheduled arrival and ends when the final request finishes or fails. Include the final drain, not just the time requests were submitted. Output throughput is successful requests’ generated tokens divided by that window. Request throughput is successful requests divided by it. Keep the number of requests with known token counts beside token throughput; a missing usage field cannot become an invented count.

Goodput counts successful requests that meet every configured SLO. A request meeting TTFT but missing TPOT fails the combined objective. Missing TPOT cannot pass a TPOT SLO; failed requests remain in the attainment denominator.

def summarize(traces, duration, ttft_slo=None, tpot_slo=None, e2e_slo=None):
    """Throughput counts successful work; goodput requires every configured SLO.
    Failed requests remain in the denominator; missing token timing cannot pass a TPOT SLO.  (Your engine: Chapter 44)
    """
    if duration <= 0 or not traces:
        raise ValueError("Need traces and a positive measurement window")
    good = [t for t in traces if t.error is None]
    met = [t for t in good if all(limit is None or value is not None and value <= limit
                                for value, limit in ((t.ttft, ttft_slo), (t.tpot, tpot_slo), (t.e2e, e2e_slo)))]
    gaps, token_gaps = [], []
    for t in good:
        intervals = [b - a for a, b in zip(t.content_times, t.content_times[1:])]
        gaps.extend(intervals)
        # Only report token ITL when the stream proves every content event is one token,
        # and there are no invisible tokens omitted from the events.
        if t.event_token_counts and all(n == 1 for n in t.event_token_counts) \
                and t.output_tokens == len(t.content_times):
            token_gaps.extend(intervals)
    measured = [t for t in good if t.output_tokens is not None]
    return {"requests": len(traces), "successful": len(good), "failed": len(traces) - len(good),
            "duration_seconds": duration, "request_throughput_per_second": len(good) / duration,
            "output_token_throughput_per_second": sum(t.output_tokens for t in measured) / duration if measured else None,
            "requests_with_token_counts": len(measured), "goodput_requests_per_second": len(met) / duration,
            "slo_attainment": len(met) / len(traces),
            "latency_seconds": {"ttft": percentiles([t.ttft for t in good]),
                                "tpot": percentiles([t.tpot for t in good]),
                                "itl": percentiles(token_gaps), "chunk_gap": percentiles(gaps),
                                "e2e": percentiles([t.e2e for t in good]),
                                "client_queue": percentiles([t.dispatched - t.scheduled for t in traces])}}

Worked arithmetic, not a performance measurement: over 2 seconds, successful requests produce 3 and 5 tokens and a third request fails. Output throughput is 4 tokens/s, request throughput is 1 request/s. If only the first request meets every SLO, goodput is 0.5 requests/s and attainment is 1/3. Reporting only successful requests would hide the failure.

Percentiles use linear interpolation between sorted samples. Run enough requests to support a tail claim: the p99 of ten samples is almost the maximum, with very little statistical evidence. Repeat trials, report trial variation, and sweep arrival rate until SLO attainment falls below the service target. Goodput usually rises, peaks and then falls as queueing overwhelms the engine.

Run a workload

Create prompts.jsonl with fixed text and output limits:

{"prompt":"Explain why a paged KV cache saves memory.","max_tokens":128}
{"prompt":"Write a short example of continuous batching.","max_tokens":128}

For a real run use hundreds or thousands of rows, with representative lengths. Keep the exact file and hash it. Test at least short-prefill/long-decode, long-prefill/short-decode, mixed lengths and deliberately shared prefixes. Distinguish a cold run (fresh process/cache and unrelated prompts) from a warm run (declared shared prefixes and cache state). Warmup also warms prefixes; do not call a workload cold just because compilation is finished.

python -m izh.bench --base-url http://127.0.0.1:8000 --model Qwen3-0.6B \
  --workload prompts.jsonl --arrival poisson --rate 4 --seed 7 \
  --max-concurrency 256 --warmup 4 --timeout 120 \
  --ttft-slo 1.0 --tpot-slo 0.05 --e2e-slo 10 \
  --metadata experiment.json --output runs/izh-poisson-4.json

python -m izh.bench --base-url http://127.0.0.1:8000 --model Qwen3-0.6B \
  --workload prompts.jsonl --arrival burst --burst-size 16 --rate 4 \
  --output runs/izh-burst-4.json

SLOs in those commands are experiment inputs, not claims about achieved latency. The client uses raw /v1/completions, temperature=0, n=1, no chat template, and each row’s max_tokens. EOS may finish early: record actual token counts. If testing fixed output lengths, configure every server’s supported EOS policy consistently and retain that configuration. Do not send an extension one engine ignores and assume it took effect.

Accuracy beside speed

First check greedy equality against an independent generator, starting with identical token IDs. Report the first divergent token. Once one token differs, later histories differ; a full-sequence logit comparison no longer diagnoses the first error. A candidate that returns immediately or uses the wrong tokenizer can look wonderfully fast.

Then compare logits on the same teacher-forced history:

def logit_quality(reference, candidate, ids, mask=None):
    """Logits [B,T,V], IDs [B,T]. Score next tokens, excluding optional padding.
    KL is KL(reference || candidate), teacher-forced on the same history.  (Your engine: Chapter 44)
    """
    if reference.shape != candidate.shape or reference.shape[:2] != ids.shape or ids.shape[1] < 2:
        raise ValueError("Quality comparison needs matching logits and at least two tokens")
    targets = ids[:, 1:]
    valid = torch.ones_like(targets, dtype=torch.bool) if mask is None else (mask[:, 1:] & mask[:, :-1]).bool()
    if not valid.any():
        raise ValueError("No next-token positions to score")
    ref = reference[:, :-1].float().log_softmax(-1)
    cand = candidate[:, :-1].float().log_softmax(-1)
    rloss = -ref.gather(-1, targets[..., None]).squeeze(-1)[valid].mean()
    closs = -cand.gather(-1, targets[..., None]).squeeze(-1)[valid].mean()
    kl = (ref.exp() * (ref - cand)).sum(-1)[valid].mean()
    return {"tokens": int(valid.sum()), "reference_nll": float(rloss), "candidate_nll": float(closs),
            "reference_perplexity": math.exp(float(rloss)), "candidate_perplexity": math.exp(float(closs)),
            "kl_reference_candidate": float(kl),
            "teacher_forced_argmax_agreement": float((ref.argmax(-1) == cand.argmax(-1))[valid].float().mean()),
            "max_abs_logit_difference": float((reference - candidate).abs().max())}


def greedy_equality(reference_generate, candidate_generate, prompts, max_tokens=16):
    """Call two independent {request_id: tokens} generators; report the first divergence."""
    ref, cand = reference_generate(prompts, max_tokens), candidate_generate(prompts, max_tokens)
    differences = []
    for rid in prompts:
        a, b = ref[rid], cand[rid]
        if a != b:
            at = next((i for i, (x, y) in enumerate(zip(a, b)) if x != y), min(len(a), len(b)))
            differences.append({"request_id": rid, "position": at,
                                "reference": a[at] if at < len(a) else None, "candidate": b[at] if at < len(b) else None})
    return {"requests": len(prompts), "equal": len(prompts) - len(differences), "differences": differences}

Next-token negative log likelihood excludes the first token and masked padding; perplexity is its exponential. KL(reference || candidate) measures how much the whole distribution changed, even if argmax did not. Use the same tokens, position conventions and mask in both models. Compare FP32 first, then deployment precision and quantization. Exact greedy equality is a strong regression test, but small floating-point differences near a tied argmax can change a trajectory without proving a major quality loss (Chapter 13).

Finally run a small task evaluation. The supplied runner accepts a local JSONL subset with question and answer fields, including GSM8K’s #### numeric answer convention:

async def evaluate_endpoint(client, base_url, model, examples, max_tokens=256):
    """A reproducible greedy, zero-shot numeric-answer smoke evaluation.  (Your engine: Chapter 44)
    Each row retains the question, expected answer, full response and HTTP failures.
    """
    records = []
    for i, row in enumerate(examples):
        expected = numeric_answer(row["answer"])
        if expected is None:
            raise ValueError(f"Example {i} has no numeric reference answer")
        prompt = f"Question: {row['question']}\nSolve step by step. End with #### followed by the numeric answer.\nAnswer:"
        result = {"index": i, "question": row["question"], "answer": row["answer"], "prompt": prompt, "correct": False}
        try:
            response = await client.post(base_url.rstrip("/") + "/v1/completions",
                                         json={"model": model, "prompt": prompt, "temperature": 0, "max_tokens": max_tokens})
            response.raise_for_status()
            result["output"] = response.json()["choices"][0]["text"]
            result["correct"] = numeric_answer(result["output"]) == expected
        except Exception as error:
            result["error"] = f"{type(error).__name__}: {error}"
        records.append(result)
    if not records:
        raise ValueError("Evaluation subset is empty")
    return {"protocol": "zero-shot numeric extraction, greedy completions; not the official GSM8K protocol",
            "model": model, "max_tokens": max_tokens, "examples": len(records),
            "accuracy": sum(r["correct"] for r in records) / len(records), "records": records}
python -m izh.evaluation --base-url http://127.0.0.1:8000 --model Qwen3-0.6B \
  --data data/gsm8k-subset.jsonl --limit 32 --max-tokens 256 --output runs/izh-eval.json

The reader supplies that file; it isn’t bundled or downloaded. The runner retains prompts, responses and failures, uses greedy zero-shot completions and extracts a numeric answer. This is a smoke protocol, not the official few-shot GSM8K evaluation. A small subset has large uncertainty and the extractor can accept a stray final number. For a published score, use the benchmark’s defined protocol and a reviewed harness. Preserve dataset revision, subset IDs and licenses with your experiment.

Launch the competitors with the same checkpoint

These recipes use one GPU, the same local Qwen3-0.6B checkpoint, FP16 weights and KV, a 4,096-token per-request limit and up to 16 active sequences. Start from one pinned HF revision and convert that directory for GGUF. Use each engine’s own supported environment; mixing all five dependency stacks into one environment is unnecessary. Record installed versions or git SHAs and actual resolved startup settings. CLI recipes were checked against the linked upstream documentation; competitor execution is recorded separately in Appendix F.

Run servers one at a time on port 8000 and query /v1/models for the model name the benchmark must send. Set CUDA_VISIBLE_DEVICES=0 in the server shell. Check prompt IDs/BOS behavior before comparing greedy outputs; identical text does not guarantee identical tokenization.

The book’s engine:

python serve.py --model-dir models/Qwen3-0.6B --served-model-name Qwen3-0.6B \
  --device cuda --dtype fp16 --max-model-len 4096 --max-num-seqs 16 \
  --max-num-batched-tokens 2048 --attention-backend triton --kv-cache-dtype auto \
  --cuda-graphs auto --port 8000

vLLM (serve arguments):

vllm serve models/Qwen3-0.6B --served-model-name Qwen3-0.6B \
  --dtype float16 --kv-cache-dtype auto --tensor-parallel-size 1 \
  --max-model-len 4096 --max-num-seqs 16 --max-num-batched-tokens 2048 --port 8000

SGLang (server arguments):

python -m sglang.launch_server --model-path models/Qwen3-0.6B \
  --served-model-name Qwen3-0.6B --dtype float16 --kv-cache-dtype auto \
  --tp-size 1 --context-length 4096 --max-running-requests 16 \
  --chunked-prefill-size 2048 --port 8000

llama.cpp (server arguments), from its checkout with a CUDA build:

python convert_hf_to_gguf.py /absolute/path/to/models/Qwen3-0.6B \
  --outfile /absolute/path/to/models/Qwen3-0.6B-F16.gguf --outtype f16
build/bin/llama-server --model /absolute/path/to/models/Qwen3-0.6B-F16.gguf \
  --alias Qwen3-0.6B --n-gpu-layers 999 --parallel 16 --kv-unified \
  --kv-unified-per-slot 4096 --batch-size 2048 --ubatch-size 512 \
  --cache-type-k f16 --cache-type-v f16 --port 8000

With these current flags, the unified pool is sized for 16 × 4,096 positions. Older revisions use different context/slot semantics; preserve the resolved per-slot context printed at startup rather than assuming --ctx-size 4096 gives every one of sixteen slots 4,096 tokens.

ExLlamaV3 through TabbyAPI (sample config), from its checkout, with config.yml:

network:
  host: 127.0.0.1
  port: 8000
  disable_auth: true
  allowed_origins: []
model:
  model_dir: /absolute/path/to/models
  model_name: Qwen3-0.6B
  backend: exllamav3
  max_seq_len: 4096
  cache_size: 65536
  cache_mode: FP16
  max_batch_size: 16
  tensor_parallel: false
python main.py

This local benchmark configuration loads the original unquantized checkpoint; ExLlamaV3’s linear implementation has an unquantized path. Its native quantized deployment deserves a separate EXL3 experiment, with bitrate and accuracy recorded, alongside separate GGUF, GPTQ/AWQ and FP8 experiments for the other engines. A 4-bit result is not the same precision experiment as the baseline.

TensorRT-LLM (CLI, LLM arguments), with trtllm-bench.yml:

dtype: float16
max_batch_size: 16
max_num_tokens: 2048
max_seq_len: 4096
tensor_parallel_size: 1
pipeline_parallel_size: 1
trtllm-serve models/Qwen3-0.6B --backend pytorch --config trtllm-bench.yml --port 8000

These match model, precision, context and sampling, but not every internal policy. Graph buckets, prefix caches, total KV allocation, prefill tiling and batch-token limits differ. Measure each engine’s normal optimized serving configuration first and disclose those differences. Then run controlled ablations (prefix cache, graphs, speculation), changing one setting at a time. Restart between cold trials; use deliberately reused prefixes for warm trials. For a memory-normalized comparison, match usable KV bytes explicitly rather than trusting unrelated default utilization fractions.

A results table that earns its numbers

There are no competitor measurements in this table. Fill a row only after saving its trace JSON, startup log, revisions, workload hash and accuracy result. A local random-model smoke run validates transport and metric plumbing; it cannot rank these engines.

engine / revisionprecision / KVworkload / rateoutput tok/sTTFT p50 / p99TPOT p50 / p99goodput req/sfailuresgreedy / KL / task score
izhFP16 / FP16not run——————
vLLMFP16 / FP16not run——————
SGLangFP16 / FP16not run——————
llama.cppF16 GGUF / F16not run——————
ExLlamaV3 / TabbyAPIunquantized / FP16not run——————
TensorRT-LLMFP16 / FP16not run——————

Keep per-token ITL and chunk-gap percentiles beside this table when studying streaming stalls. Record CPU usage and GPU memory alongside throughput: moving a bottleneck into a tokenizer process or exhausting host memory is still a system bottleneck.

Harden what you have measured

The service now has a raw-body limit before JSON parsing, a body-read deadline, a total generation deadline, bounded pending requests and a bounded number of buffered stream updates. If a consumer falls behind, its request is aborted and its buffer becomes an error, releasing its KV and adapter lease. It does not let one stalled connection retain memory indefinitely.

AsyncLLM.stop marks the service draining, refuses new generation, gives existing consumers 30 seconds to finish, then aborts remaining requests, shuts down and joins the core worker. A process still alive after the join deadline is terminated; a Python thread cannot be forcibly stopped safely. /health reports fatal engine failure; /ready additionally requires loaded model metadata and refuses readiness while draining. The launcher emits structured request metadata (ID, path, status and elapsed time) without logging prompts. Stream failures after HTTP 200 still need stream/error counters; the HTTP status alone isn’t sufficient.

Deployment checklist:

failure to exercisebehavior to verifysupplied mechanism / remaining integration
slow or oversized bodyreject before expensive parsing/allocationRequestLimits: byte cap and body deadline, including chunked bodies
overloaded admissionbounded queues; clear retryable errormax_pending, 503, scheduler context/KV limits; add per-tenant quotas
long or stuck generationcancel at a total deadlinerequest_timeout, finally abort; test stream error delivery through your proxy
stalled/disconnected consumerrelease KV and adapter leasesupdate cap, cancellation, core abort; test proxy disconnect propagation
deployment shutdownstop admission, drain, then bounded abort/joinAsyncLLM.stop; align orchestrator/proxy grace periods and stop routing before teardown
failed model load / core crashreadiness fails; all affected clients failready/fatal messages, /health, /ready; exercise process death and restart
malformed request or unknown adapterexplicit client errorschema/parameter validation; monitor 4xx and avoid retry storms
incident investigationcorrelate request/engine events without text leakagerequest IDs, structured metadata, Prometheus counters; connect to your log collector
checkpoint or adapter replacementno stale KV or mid-request weight changedrain base-model updates; adapter leases and content keys

Rate limiting, identity, rolling-deployment routing, persistent observability, host-level limits and recovery from device failure belong to the deployment around the engine. Test them with faults, not just a successful curl. Chapter 45 distinguishes the implemented mechanisms from the work still needed for an operational service.

Build it

Engine milestone 44: measure the service. Implement arrival_offsets, request and summarize in engine/bench.py, and logit_quality and evaluate_endpoint in engine/evaluation.py. The CLI entry points, trace properties, SSE parser and greedy-equality helper are provided.

pytest tests/test_ch44_benchmarking.py
python run.py benchmark
python -m izh.bench --help
python -m izh.evaluation --help

The tests use fixed clocks for denominators/SLOs, fragmented UTF-8 SSE streams for parsing, explicit HTTP and truncated-stream failures, seeded arrival sequences, independent cross-entropy calculations, first-divergence checks, a task fixture with a failed request, body limits and slow-consumer cancellation. Their timing fixtures are not benchmark measurements.

Stretch exercises

  1. ★★ Sweep offered rate, plot p99 TTFT and goodput, and repeat three trials at each point. Find the knee where waiting-queue growth begins. Where: terminal: the benchmark CLI recipes above; save trial results and plots in experiments/.
  2. ★★ Add separate cold/warm prefix workloads and compare equal memory budgets. How much of an observed gain comes from cache hits rather than faster kernels? Where: workload construction in engine/bench.py or experiments/ch44.py (create it) using its benchmark function.
  3. ★★★ Add a chat or multimodal workload with the same rendered prompt/processor in every engine. Record image-encoder time separately from decoder prefill. Where: request construction in engine/bench.py, with chat/image preparation shared by experiments/ch44.py (create it).
  4. ★★ Inject core-process failure and proxy disconnects during a burst. Verify bounded client failure time, memory release and readiness transitions. Where: new experiments/service_failures.py, driving the server as a subprocess and an interruptible proxy.
  5. ★★★ Compare INT4 and FP8 variants with teacher-forced KL and a full task protocol. Report a speed-quality-memory curve rather than one winning number. Where: experiments/ch44.py (create it), using engine.evaluation.logit_quality and evaluate_endpoint alongside engine.bench.benchmark.

Check your understanding

  1. Why can a closed-loop client conceal overload?
  2. Which timestamp changes when the client waits for a connection, and which should remain fixed?
  3. Why can’t a five-token SSE chunk supply five real ITL measurements?
  4. Why is TPOT undefined for a one-token response?
  5. Why must failed requests remain in SLO attainment’s denominator?
  6. Why compare logits on a teacher-forced history after greedy outputs diverge?
  7. Why are a 4-bit model and an FP16 model different benchmark experiments?
  8. What should readiness report while a process is healthy but draining?

Going deeper