Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

35. Text in, text out: tokenizers, templates and tool calls

In this chapter

  • Loading any byte-level or SentencePiece-style BPE tokenizer from tokenizer.json, matching Hugging Face's tokenizers token for token, with no dependency on it.
  • Streaming text from a stream of tokens without ever printing half a character.
  • Stop strings that span tokens, and holding back text that might be the start of one.
  • Chat templates rendered exactly as transformers renders them, safely.
  • Turning a model's tool calls and reasoning into the API's structured fields, while streaming, and forcing a tool call with Chapter 34's constrained decoding.

You will build

Tokenizer.encode and Tokenizer.bpe, IncrementalDetokenizer and StopChecker (engine/serve/tokenizer.py); render_chat and TagStreamer.feed (engine/serve/chat.py).

Time: 5-7 hours. GPU: not needed.

The text boundary is part of the engine

Chapter 18’s engine used Hugging Face’s tokenizer “at the text boundary” and stopped there. A server can’t. Every request crosses the boundary twice: its messages become a prompt through a chat template and a tokenizer, and its tokens become streamed text, stop-string checks, tool calls and reasoning. Mistakes here look like model bugs: a chat template missing one newline costs measurable accuracy, a detokenizer that prints each token’s bytes shows � in every emoji, a stop string split across two tokens is never noticed, and a tool call printed as text breaks every agent that called the API.

Owning this code also matters for speed and deployment. The engine of Part VIII has no dependency on transformers at serving time, and Chapter 36 runs tokenization in a separate process from the engine loop, so it must be code you control.

Loading tokenizer.json

A tokenizer.json file describes a pipeline: a normalizer (often Unicode NFC), a pre-tokenizer that splits text into words, a model (BPE: a vocabulary and a ranked list of merges), added tokens that are matched before anything else, and a decoder. Two families of BPE cover nearly all open models:

byte-level BPESentencePiece-style BPE
modelsGPT-2, Llama 3, Qwen, DeepSeek, Mistral (Tekken)Llama 2, Mistral v0.1-v0.3, Gemma
pre-tokenizera regex splits words; each word’s UTF-8 bytes are mapped to printable charactersspaces become ▁; words start at ▁
unknown textimpossible: all 256 bytes are in the vocabularycharacters outside the vocabulary fall back to <0x00>-<0xFF> byte tokens
decodermap characters back to bytes, UTF-8 decode▁ → space, byte tokens → bytes, drop the leading space

Byte-level vocabularies store tokens as strings over a 256-character alphabet that GPT-2 introduced, so that every byte has a printable stand-in (Ġ is a space, Ċ a newline):

@lru_cache(maxsize=1)
def bytes_to_unicode():
    """GPT-2's reversible map from the 256 byte values to printable characters: printable
    Latin-1 bytes map to themselves, the rest to code points 256 and up. Byte-level vocabularies
    store tokens as strings over this alphabet ("Ġ" is a space, byte 0x20)."""
    keep = list(range(ord("!"), ord("~") + 1)) + list(range(ord("¡"), ord("¬") + 1)) + list(range(ord("®"), ord("ÿ") + 1))
    chars, extra = {}, 0
    for b in range(256):
        if b in keep:
            chars[b] = chr(b)
        else:
            chars[b] = chr(256 + extra)
            extra += 1
    return chars

Encoding runs the pipeline. Added tokens such as <|im_start|> are found first, longest first, so a chat’s control tokens are never split or merged with text. The rest is normalized, split into words by the pre-tokenizer’s regex, and each word is merged independently:

class Tokenizer:
    def __init__(self, spec, config=None):
        model = spec["model"]
        if model.get("type", "BPE") != "BPE":
            raise ValueError(f"Only BPE tokenizers are implemented, not {model.get('type')}")
        self.vocab = dict(model["vocab"])
        self.ranks = {pair: i for i, pair in enumerate(_merges(model.get("merges", [])))}
        self.byte_fallback = model.get("byte_fallback", False)
        self.ignore_merges = model.get("ignore_merges", False)
        self.unk = model.get("unk_token")
        self.added = {t["content"]: t for t in spec.get("added_tokens", [])}
        for content, token in self.added.items():
            self.vocab.setdefault(content, token["id"])
        self.id_to_token = {i: t for t, i in self.vocab.items()}
        self.special_ids = {t["id"] for t in self.added.values() if t.get("special")}
        pre = _flatten(spec.get("pre_tokenizer"), "pre")
        self.splits = [regex.compile(p["pattern"].get("Regex") or regex.escape(p["pattern"]["String"]))
                       for p in pre if p["type"] == "Split"]
        byte_level = [p for p in pre if p["type"] == "ByteLevel"]
        self.byte_level = bool(byte_level) or any(d["type"] == "ByteLevel" for d in _flatten(spec.get("decoder"), "dec"))
        if byte_level and byte_level[0].get("use_regex", True) and not self.splits:
            self.splits = [regex.compile(r"'s|'t|'re|'ve|'m|'ll|'d| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+(?!\S)|\s+")]
        self.add_prefix_space = bool(byte_level and byte_level[0].get("add_prefix_space"))
        metaspace = [p for p in pre if p["type"] == "Metaspace"]
        if metaspace and metaspace[0].get("split", True):
            self.splits = self.splits + [regex.compile("▁?[^▁]+|▁+(?=▁)|▁+$")]     # words start at ▁
        norms = _flatten(spec.get("normalizer"), "norm")
        self.nfc = any(n["type"] == "NFC" for n in norms)
        self.space_to_meta = bool(metaspace) or any(n["type"] == "Replace" and n["content"] == "▁" for n in norms)
        scheme = metaspace[0].get("prepend_scheme", "always") if metaspace else "always"
        self.prepend_meta = self.space_to_meta and (scheme != "never" if metaspace else
                                                   any(n["type"] == "Prepend" for n in norms))
        decoders = _flatten(spec.get("decoder"), "dec")
        self.strip_leading_space = any(d["type"] == "Strip" and d.get("start", 0) for d in decoders) or \
            any(d["type"] == "Metaspace" and d.get("prepend_scheme", "always") != "never" for d in decoders)
        contents = sorted(self.added, key=len, reverse=True)
        self.added_pattern = regex.compile("|".join(regex.escape(c) for c in contents)) if contents else None
        self.byte_map = bytes_to_unicode()
        self.byte_unmap = {c: b for b, c in self.byte_map.items()}
        config = config or {}
        self.chat_template = config.get("chat_template")
        self.bos_token = _content(config.get("bos_token"))
        self.eos_token = _content(config.get("eos_token"))
        self.cache = {}

    @classmethod
    def from_dir(cls, directory):
        directory = Path(directory)
        config_path = directory / "tokenizer_config.json"
        config = json.loads(config_path.read_text()) if config_path.exists() else {}
        return cls(json.loads((directory / "tokenizer.json").read_text()), config)

    @property
    def vocab_size(self):
        return max(self.id_to_token) + 1

    def token_to_id(self, token):
        return self.vocab.get(token)

    @property
    def eos_token_id(self):
        return self.vocab.get(self.eos_token) if self.eos_token else None

    # ------------------------------------------------------------ encoding
    def encode(self, text):
        """Text -> token ids: added tokens first, then normalize, pre-tokenize and merge.  (Your engine: Chapter 35)"""
        ids, start = [], 0
        for match in (self.added_pattern.finditer(text) if self.added_pattern else ()):
            ids.extend(self._encode_ordinary(text[start:match.start()]))
            ids.append(self.vocab[match.group()])
            start = match.end()
        ids.extend(self._encode_ordinary(text[start:]))
        return ids

    def _encode_ordinary(self, text):
        if not text:
            return []
        if self.nfc:
            text = unicodedata.normalize("NFC", text)
        if self.byte_level:
            if self.add_prefix_space and not text.startswith(" "):
                text = " " + text
            ids = []
            for word in self.pre_tokenize(text):
                ids.extend(self.bpe("".join(self.byte_map[b] for b in word.encode("utf-8"))))
            return ids
        if self.space_to_meta:
            text = text.replace(" ", "▁")
            if self.prepend_meta and not text.startswith("▁"):
                text = "▁" + text
        ids = []
        for word in (self.pre_tokenize(text) if self.splits else [text]):
            ids.extend(self.bpe(word))
        return ids

    def pre_tokenize(self, text):
        """Split into pieces with each Split regex, keeping both matches and gaps ("Isolated")."""
        pieces = [text]
        for pattern in self.splits:
            out = []
            for piece in pieces:
                pos = 0
                for m in pattern.finditer(piece):
                    if m.start() > pos:
                        out.append(piece[pos:m.start()])
                    if m.end() > m.start():
                        out.append(m.group())
                    pos = m.end()
                if pos < len(piece):
                    out.append(piece[pos:])
            pieces = out
        return pieces

    def bpe(self, word):
        """Merge the lowest-ranked adjacent pair until none is in the merge table.  (Your engine: Chapter 35)"""
        if word in self.cache:
            return self.cache[word]
        if self.ignore_merges and word in self.vocab:
            return [self.vocab[word]]
        parts = list(word)
        while len(parts) > 1:
            best, where = None, -1
            for i, pair in enumerate(zip(parts, parts[1:])):
                rank = self.ranks.get(pair)
                if rank is not None and (best is None or rank < best):
                    best, where = rank, i
            if best is None:
                break
            parts[where:where + 2] = [parts[where] + parts[where + 1]]
        ids = []
        for part in parts:
            if part in self.vocab:
                ids.append(self.vocab[part])
            elif self.byte_fallback:
                ids.extend(self.vocab[f"<0x{b:02X}>"] for b in part.encode("utf-8"))
            elif self.unk is not None:
                ids.append(self.vocab[self.unk])
            else:
                raise ValueError(f"Cannot encode {part!r}: not in the vocabulary and no byte fallback")
        if len(self.cache) < 200_000:
            self.cache[word] = ids
        return ids

    # ------------------------------------------------------------ decoding
    def token_bytes(self, token_id, skip_special_tokens=False):
        """The bytes one token contributes to the text (b"" for skipped specials)."""
        token = self.id_to_token.get(token_id)
        if token is None or (skip_special_tokens and token_id in self.special_ids):
            return b""
        if token in self.added:
            return token.encode("utf-8")
        if self.byte_level:
            return bytes(self.byte_unmap[c] for c in token)
        if self.byte_fallback and len(token) == 6 and token.startswith("<0x") and token.endswith(">"):
            return bytes([int(token[3:5], 16)])
        return token.replace("▁", " ").encode("utf-8")

    def decode(self, ids, skip_special_tokens=False):
        data = b"".join(self.token_bytes(i, skip_special_tokens) for i in ids)
        text = data.decode("utf-8", errors="replace")
        if self.strip_leading_space and text.startswith(" "):
            text = text[1:]
        return text

    def vocab_bytes(self):
        """bytes per token id for constrained decoding (Chapter 34); special tokens are b"" (never text)."""
        return [b"" if i in self.special_ids else self.token_bytes(i) for i in range(self.vocab_size)]

The pre-tokenizer regex decides which merges are even possible, so it’s part of the model’s definition, not an implementation detail. Qwen’s splits letters from digits and every digit from the next (\p{N}); Llama 3’s allows digit groups of up to three (\p{N}{1,3}); GPT-2’s attaches a leading space to words. These patterns use Unicode property classes (\p{L} is any letter in any script), which Python’s built-in re lacks, so the loader uses the regex package.

bpe is Chapter 4’s algorithm with a learned merge table: repeatedly merge the adjacent pair with the lowest rank. Two production details are worth noticing. Results are cached per word, because natural text repeats words constantly; that cache is why this pure-Python implementation keeps up with Rust on ordinary text. And Llama 3’s ignore_merges flag says that a word that’s already in the vocabulary is one token, without running the merges.

The test trains three tokenizers with Hugging Face’s tokenizers library: a Qwen-style byte-level one with its split regex and NFC, a GPT-2-style one, and a SentencePiece-style one with byte fallback. It then checks that this loader produces exactly the same token IDs and the same decoded text on sixteen strings: control tokens, whitespace runs, contractions, code, Chinese, Japanese, emoji with zero-width joiners, and characters the tokenizer never saw.

vocab_bytes() gives each token’s bytes, which is what Chapter 34’s constrained decoding walks. Special tokens map to b"", so a pattern can never be satisfied by <|im_end|>. Tokens that were added but aren’t special, like Qwen3’s <think>, are ordinary text.

Streaming text without half characters

The engine produces tokens one at a time, and the server must stream text. Decoding each token’s bytes separately fails whenever a token ends inside a UTF-8 character, which byte-level BPE does routinely for rare characters:

tokens:            a | ␠ | F0 9F | A7 91 | E2 80 8D | ...
per-token decode:  "a ��������� on ����������������"
incremental:       "a", " ", "", "🧑", "", "", "‍", "", "🚀", " on", ...

The fix is an incremental UTF-8 decoder: it returns every complete character and holds back an incomplete one until the bytes that finish it arrive. Python’s codecs module has one:

class IncrementalDetokenizer:
    """Turns a stream of token ids into a stream of text deltas.  (Your engine: Chapter 35)

    A token can end in the middle of a UTF-8 character (byte-level BPE splits "é" or an emoji
    across tokens), so printing each token's bytes as they arrive would print replacement
    characters. Bytes go through an incremental UTF-8 decoder, which holds back an incomplete
    trailing character until the token that completes it arrives.
    """

    def __init__(self, tokenizer, skip_special_tokens=True):
        self.tokenizer, self.skip = tokenizer, skip_special_tokens
        self.decoder = codecs.getincrementaldecoder("utf-8")(errors="replace")
        self.text = ""

    def add(self, token_ids):
        data = b"".join(self.tokenizer.token_bytes(t, self.skip) for t in token_ids)
        delta = self.decoder.decode(data, final=False)
        self.text += delta
        return delta

    def flush(self):
        delta = self.decoder.decode(b"", final=True)
        self.text += delta
        return delta

Engines that call a string-level decode (vLLM, Text Generation Inference) get the same effect with a more complex trick: they decode a window of recent tokens twice, with and without the newest one, and stream the difference only if it doesn’t end in a replacement character. Working at the byte level makes the problem disappear, because bytes are what tokens really are.

Stop strings

stop=["</answer>"] must end generation when the text contains </answer>, even if it arrives as </ans + wer>, and the client must never see the stop string, or any part of it, unless include_stop_str_in_output is set. Two rules handle that:

  1. Search the text the new delta could have completed: the last len(longest stop) - 1 characters before it, plus the delta.
  2. Hold back any tail of the text that is a prefix of a stop string, until the next delta proves it isn’t. "2.</ans" streams "2." and keeps "</ans". If the next delta is "wer>", the stop is complete and nothing more is sent; if it’s "ia", the held text is released as "</ansia".
class StopChecker:
    """Stop strings on streamed text.  (Your engine: Chapter 35)

    A stop string can span several deltas, so the checker searches the text that the new delta
    could have completed, and it holds back (doesn't stream yet) any tail of the text that is
    the start of a stop string, so that a client never sees "</ans" before the stop is found.
    """

    def __init__(self, stops, include_stop=False):
        self.stops, self.include = [s for s in stops if s], include_stop
        self.longest = max((len(s) for s in self.stops), default=0)
        self.text, self.sent = "", 0
        self.stopped = None

    def add(self, delta):
        """Append delta; returns the text that may be streamed now. self.stopped is set on a match."""
        old = len(self.text)
        self.text += delta
        if self.stops:
            window = max(0, old - self.longest + 1)
            hits = [(self.text.find(s, window), s) for s in self.stops]
            hits = [(i, s) for i, s in hits if i >= 0]
            if hits:
                index, stop = min(hits)
                self.text = self.text[:index + (len(stop) if self.include else 0)]
                self.stopped = stop
                return self._release(len(self.text))
        hold = 0
        for s in self.stops:                              # the longest tail that begins some stop string
            for k in range(min(len(s) - 1, len(self.text)), 0, -1):
                if self.text.endswith(s[:k]):
                    hold = max(hold, k)
                    break
        return self._release(len(self.text) - hold)

    def finish(self):
        return self._release(len(self.text))

    def _release(self, upto):
        out = self.text[self.sent:upto]
        self.sent = max(self.sent, upto)
        return out

Stop strings are checked on detokenized text, so they belong to the frontend (Chapter 36), not to the engine core. When the frontend finds one, it aborts the request in the core, which frees its blocks.

Chat templates

A chat model saw every training conversation rendered in one exact format. Qwen uses ChatML:

<|im_start|>system
You are terse.<|im_end|>
<|im_start|>user
Weather in Paris?<|im_end|>
<|im_start|>assistant

The format is defined by a Jinja template shipped in the model’s tokenizer_config.json, and the model’s quality depends on rendering it exactly: the newlines, the tools block, whether a <think> block is pre-filled. The only reliable way to match is to render the template the way transformers does:

def render_chat(template, messages, tools=None, add_generation_prompt=True, **variables):
    """Render a Hugging Face chat template exactly as transformers' apply_chat_template does.  (Your engine: Chapter 35)

    Same environment: a sandbox (templates come from downloaded files and must not run code),
    trim_blocks and lstrip_blocks, the loopcontrols extension, and the helpers templates use:
    raise_exception, strftime_now and a tojson filter that keeps non-ASCII text.
    """
    from jinja2.ext import loopcontrols
    from jinja2.sandbox import ImmutableSandboxedEnvironment

    def raise_exception(message):
        raise ValueError(message)

    def tojson(value, ensure_ascii=False, indent=None, separators=None, sort_keys=False):
        return json.dumps(value, ensure_ascii=ensure_ascii, indent=indent, separators=separators, sort_keys=sort_keys)

    env = ImmutableSandboxedEnvironment(trim_blocks=True, lstrip_blocks=True, extensions=[loopcontrols])
    env.filters["tojson"] = tojson
    env.globals["raise_exception"] = raise_exception
    env.globals["strftime_now"] = lambda fmt: datetime.now().strftime(fmt)
    return env.from_string(template).render(messages=messages, tools=tools,
                                            add_generation_prompt=add_generation_prompt, **variables)

Three details make it match. Whitespace control: trim_blocks and lstrip_blocks remove the newlines and indentation around {% %} tags, which templates are written to expect. Helpers: templates call raise_exception, strftime_now and a tojson filter that, unlike Jinja’s built-in one, keeps non-ASCII characters and accepts indent. Loop controls: templates use {% continue %}, which needs Jinja’s loopcontrols extension. And one detail makes it safe: the sandbox. A template is code that came with a downloaded model; the sandboxed environment refuses attribute access that could reach Python internals, so a malicious template can’t run code on the server.

The book ships CHATML_TEMPLATE, a ChatML template with tools and switchable thinking in the style of Qwen3’s, for models without one (the test models). The test renders a conversation with a system prompt, a tool call, a tool result and non-ASCII text, with and without tools and with thinking on and off, and checks that the result equals transformers’ apply_chat_template character for character.

Tool calls and reasoning

Tool calling is a convention on top of text. The template lists the tools in the system prompt; a model trained for it answers with

<tool_call>
{"name": "get_weather", "arguments": {"city": "Paris"}}
</tool_call>

and the API must return it as structured data: "tool_calls": [{"id": ..., "type": "function", "function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}}] (the arguments are a JSON string in OpenAI’s format). Reasoning models similarly write <think>...</think> before the answer, which the API returns as reasoning_content.

Parsing a complete answer is a regex. Parsing a stream is harder: the text arrives a few characters at a time, <tool_call> can be split across deltas, and the client must never see half a tag as content. TagStreamer splits a stream into “outside” and “inside” segments for one pair of tags, holding back any tail that might start a tag:

TOOL_OPEN, TOOL_CLOSE = "<tool_call>", "</tool_call>"


def parse_tool_calls(text):
    """Complete output -> (content, [tool calls in OpenAI's format]). Hermes/Qwen style:
    <tool_call>{"name": ..., "arguments": {...}}</tool_call>, possibly several."""
    calls = []
    for body in re.findall(re.escape(TOOL_OPEN) + r"(.*?)" + re.escape(TOOL_CLOSE), text, re.S):
        calls.append(openai_tool_call(json.loads(body)))
    content = re.sub(re.escape(TOOL_OPEN) + r".*?" + re.escape(TOOL_CLOSE), "", text, flags=re.S).strip()
    return content, calls


def openai_tool_call(call):
    arguments = call.get("arguments", call.get("parameters", {}))
    return {"id": f"call_{uuid.uuid4().hex[:24]}", "type": "function",
            "function": {"name": call["name"],
                         "arguments": arguments if isinstance(arguments, str) else json.dumps(arguments, ensure_ascii=False)}}


class TagStreamer:
    """Splits a text stream into "outside" and "inside" segments for one pair of tags, holding back
    any tail that might be the start of a tag.  (Your engine: Chapter 35)

    feed(delta) returns [(inside, text), ...] for text that is certain; an inside segment is
    reported once, whole, when its closing tag arrives.
    """

    def __init__(self, open_tag, close_tag, start_inside=False):
        self.open, self.close = open_tag, close_tag
        self.inside, self.buffer = start_inside, ""

    def feed(self, delta, final=False):
        self.buffer += delta
        out = []
        while True:
            tag = self.close if self.inside else self.open
            index = self.buffer.find(tag)
            if index >= 0:
                if index or self.inside:
                    out.append((self.inside, self.buffer[:index]))
                self.buffer = self.buffer[index + len(tag):]
                self.inside = not self.inside
                continue
            if self.inside and not final:
                return out                                     # wait for the closing tag
            hold = 0 if final else max((k for k in range(1, len(tag)) if self.buffer.endswith(tag[:k])), default=0)
            if len(self.buffer) > hold:
                out.append((self.inside, self.buffer[:len(self.buffer) - hold]))
                self.buffer = self.buffer[len(self.buffer) - hold:]
            return out

OutputParser chains two of them, one for <think> and one for <tool_call>, and emits OpenAI-style deltas:

class OutputParser:
    """Streams a Qwen3-style answer into OpenAI deltas: reasoning_content, content and tool_calls.

    Newlines around the reasoning are formatting, not content: "<think>\n...\n</think>\n\n"
    yields the reasoning without them and content that starts at the first real character.
    A tool call that isn't valid JSON (or is cut off) is returned as content, not dropped.
    """

    def __init__(self, reasoning=False, tools=False, starts_in_reasoning=False):
        self.reasoning = TagStreamer("<think>", "</think>", starts_in_reasoning) if reasoning else None
        self.tools = TagStreamer(TOOL_OPEN, TOOL_CLOSE) if tools else None
        self.num_calls, self.content_started = 0, False

    def feed(self, text, final=False):
        deltas = []
        segments = self.reasoning.feed(text, final) if self.reasoning else [(False, text)]
        for in_think, part in segments:
            if in_think:
                if part.strip("\n"):
                    deltas.append({"reasoning_content": part.strip("\n")})
                continue
            for in_call, piece in (self.tools.feed(part, final) if self.tools else [(False, part)]):
                if in_call:
                    try:
                        call = openai_tool_call(json.loads(piece))
                    except (ValueError, KeyError, TypeError):
                        deltas.append({"content": TOOL_OPEN + piece})
                        continue
                    deltas.append({"tool_calls": [{"index": self.num_calls, **call}]})
                    self.num_calls += 1
                    continue
                if not self.content_started:
                    piece = piece.lstrip("\n")
                    self.content_started = bool(piece)
                if piece:
                    deltas.append({"content": piece})
        return deltas

A tool call is emitted whole when its closing tag arrives. vLLM streams the arguments’ JSON incrementally instead, which a few clients rely on; that’s stretch exercise 2. A malformed or unterminated call is returned as content rather than dropped, so a client can at least see what the model said.

Forcing a tool call

tool_choice="required" (call some tool) and tool_choice={"function": {"name": "get_weather"}} (call this tool) must be guarantees, not hopes. Chapter 34’s constrained decoding makes them so: the wrapper is literal text and the arguments follow the tool’s own JSON Schema, so the regex for “call get_weather” is the escaped prefix, the schema’s regex, and the escaped suffix:

def tool_call_pattern(tools, name=None):
    """A regex that forces the model to call a tool (tool_choice="required") or one named tool:
    the call's wrapper is literal text and its arguments follow the tool's JSON Schema
    (Chapter 34)."""
    from .structured import json_schema_to_regex, regex_escape
    options = []
    for tool in tools:
        fn = tool["function"]
        if name is not None and fn["name"] != name:
            continue
        arguments = json_schema_to_regex(fn.get("parameters") or {"type": "object", "properties": {}})
        options.append(regex_escape(f'{TOOL_OPEN}\n{{"name": "{fn["name"]}", "arguments": ') + arguments
                       + regex_escape(f"}}\n{TOOL_CLOSE}"))
    if not options:
        raise ValueError(f"No tool named {name!r}")
    return "(?:" + "|".join(options) + ")"

With it, even a model that would rather chat produces a parseable call with every required argument present and of the right type.

Run it

python run.py text                    # or --model-dir to use a real checkpoint's tokenizer

Without a checkpoint, the command trains a 4,000-token Qwen-style tokenizer on The Verdict with Hugging Face’s tokenizers, then loads its tokenizer.json with the book’s loader:

{"text": "The Verdict x4 (repetitive)", "tokenizer": "izh (Python)", "tokens": 18152, "tokens_per_s": 586506}
{"text": "The Verdict x4 (repetitive)", "tokenizer": "tokenizers (Rust)", "tokens": 18152, "tokens_per_s": 480221}
{"text": "The Verdict x4 (repetitive)", "identical_ids": true}
{"text": "20,000 random words (no cache hits)", "tokenizer": "izh (Python)", "tokens": 121540, "tokens_per_s": 1009295}
{"text": "20,000 random words (no cache hits)", "tokenizer": "tokenizers (Rust)", "tokens": 121540, "tokens_per_s": 1338480}
{"text": "20,000 random words (no cache hits)", "identical_ids": true}
{"tokens": 27, "per_token_decode": "a ��������� on ����������������", "incremental_deltas": ["a", " ", "", "🧑", "", "", "‍", "", "🚀", " on", " ", "", "", "", "𝔐", "", "", "", "𝔞", "", "", "", "𝔯", "", "", "", "𝔰"]}
{"stop_streamed": ["The answer is 4", "2", ""], "final": "The answer is 42"}
{"chat_prompt_tokens": 270, "chat_prompt_tail": "n Paris?<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"}
{"parsed": [{"reasoning_content": "Need the tool."}, {"tool_calls": {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}}]}

The IDs are identical to the Rust library’s. The speeds are close on this small vocabulary, and the word cache even puts Python ahead on repetitive text; with a real 150,000-token vocabulary each uncached word takes more merge rounds, and the Rust library also offers parallel batch encoding. The number that matters more for serving is that tokenization runs on the CPU, in Python, competing with the engine loop for the interpreter, which is why Chapter 36 moves it to its own process. The astronaut emoji, a zero-width-joiner sequence the tokenizer never saw, arrives as 27 tokens, and the incremental detokenizer turns them into whole characters with empty deltas in between, where per-token decoding prints replacement characters.

Build it

Engine milestone 35: the text boundary. Implement Tokenizer.encode, Tokenizer.bpe, IncrementalDetokenizer.add, IncrementalDetokenizer.flush and StopChecker.add in engine/serve/tokenizer.py, and render_chat and TagStreamer.feed in engine/serve/chat.py (the tokenizer.json parsing, byte map, decoding, the output parser and the forced-tool-call pattern are provided).

pytest tests/test_ch35_text.py
python run.py text --impl engine

The tests check encoding and decoding against Hugging Face’s tokenizers for Qwen-style, GPT-2-style and SentencePiece-style tokenizers on sixteen hard strings; that streaming never prints a replacement character while some tokens split characters; stop strings across deltas, with and without the stop text, and false alarms; vocab_bytes for constrained decoding; chat templates against transformers with tools, tool results and thinking on and off; tool-call and reasoning parsing when streamed one character at a time, with no partial tag leaking; and the forced-tool-call pattern.

Stretch exercises

  1. ★ Load a real checkpoint’s tokenizer (Qwen3, Llama 3, Mistral) with run.py text --model-dir, and compare IDs with tokenizers on a few megabytes of mixed text. Find and fix any difference. Where: experiments/ch35.py (create it) for parity; fix differences in engine/serve/tokenizer.py.
  2. ★★ Stream tool-call arguments incrementally: emit the name as soon as "name": "..." is complete, then the arguments’ JSON text as it arrives, in OpenAI’s delta format. Where: OutputParser in engine/serve/chat.py and delta formatting in Server.chat_chunks in engine/serve/api.py.
  3. ★★ Make bpe faster for long words: a linked list of symbols and a heap of candidate merges keyed by rank and position. Measure on a 50,000-character word without spaces (some languages and code produce these). Where: Tokenizer.bpe in engine/serve/tokenizer.py.
  4. ★★★ Implement the Unigram model (T5, ALBERT, some multilingual models): Viterbi segmentation over the vocabulary’s log-probabilities. Match tokenizers on a trained Unigram tokenizer. Where: model loading in Tokenizer.from_dir and segmentation in engine/serve/tokenizer.py.

Check your understanding

  1. Why do byte-level vocabularies map bytes to printable characters rather than storing raw bytes?
  2. Why are added tokens matched before pre-tokenization, and what would go wrong otherwise?
  3. Why does the pre-tokenizer regex belong to the model’s definition?
  4. Why does per-token decoding print replacement characters, and how does an incremental UTF-8 decoder avoid them?
  5. A stop string is "STOP" and the text so far ends with "ST". What does the checker stream, and when does it release the "ST"?
  6. Why must chat templates be rendered in a sandbox?
  7. How does constrained decoding turn tool_choice from a hint into a guarantee?

Going deeper

  • Hugging Face tokenizers documentation (the pipeline: normalizers, pre-tokenizers, models, post-processors, decoders) and its tokenizer.json format; Sennrich et al., Neural Machine Translation of Rare Words with Subword Units (2016) for BPE; Kudo and Richardson, SentencePiece (2018).
  • Andrej Karpathy, Let’s build the GPT Tokenizer (2024) and his minbpe repository, for byte-level BPE and its regex pre-tokenizers.
  • Chat templates in the transformers documentation, and vLLM’s vllm/entrypoints/chat_utils.py, vllm/entrypoints/openai/tool_parsers/ and reasoning/.
  • OpenAI’s API reference for chat completions, tool calls and streaming deltas, which defines the format the parser emits.