skip to content

How does non-deterministic JSON serialization of tool definitions cause cache misses?

level: seniorimportance: should knowfreq 42%

answer

  1. same meaning, different bytes
  2. map iteration order is not a contract
  3. sort keys, freeze the tool order
  4. pin separators and escaping too
  5. assert the prefix hash in CI

basics

~20 s

A prompt cache matches bytes, not meaning. If tool schemas are serialized from a map with unstable key or element order, two logically identical prompts produce different token sequences, so one logical prefix splinters into many distinct entries that rarely match.

solid answer

~60 s

Caching compares the serialized prompt token by token, so two structurally identical tool definitions that serialize to different byte strings are two different prefixes. The usual culprits are ordering and formatting: a schema built from a hash map whose iteration order varies between processes or runtimes, a tool list assembled from a set, keys emitted in insertion order that depends on which code path registered a field first, float or unicode formatting that differs by library version or locale, and whitespace or indentation flags that differ between services. The symptom is distinctive: hit rate that is neither zero nor healthy, and that varies by pod, by process restart or by which service sent the request. The fix is canonicalization — serialize the reusable region through one pinned function that sorts object keys, imposes an explicit deterministic order on the tool array, fixes separators and escaping, and produces the identical string for identical inputs. Then hash that string, log the hash, and assert in CI that it is stable across two independently constructed requests.

code

python · 25 lines
python
import hashlib
import json

def canonical_prefix(system_prompt, tools):
    payload = {
        "system": system_prompt,
        "tools": sorted(tools, key=lambda t: t["name"]),
    }
    blob = json.dumps(
        payload, sort_keys=True, separators=(",", ":"), ensure_ascii=False
    )
    return blob, hashlib.sha256(blob.encode("utf-8")).hexdigest()[:12]

tools_a = [
    {"name": "search", "params": {"q": "string", "limit": "int"}},
    {"name": "fetch", "params": {"url": "string"}},
]
tools_b = [
    {"name": "fetch", "params": {"url": "string"}},
    {"name": "search", "params": {"limit": "int", "q": "string"}},
]

assert canonical_prefix("You are a support agent.", tools_a)[1] == \
       canonical_prefix("You are a support agent.", tools_b)[1]
print(canonical_prefix("You are a support agent.", tools_a)[1])

go deeper

for a junior

Know that the cache compares the exact serialized text, so reordered JSON keys or a differently formatted number count as a different prompt even though the meaning is unchanged.

for a middle

Explain the concrete sources — map iteration order, set-backed tool registries, separator and escaping flags, float formatting — and describe canonical serialization with sorted keys and a pinned tool order as the fix.

for a senior

Recognise the diagnostic signature: partial, host-dependent hit rate with small bounded prefix-hash cardinality, rather than the flat zero a per-request value produces. Drive it to a single shared builder and a hash logged on every call.

for a principal

Treat prefix serialization as a contract the whole platform depends on: one owned builder, CI assertions on hash stability, and a policy that dependency upgrades touching schema generation are prompt-affecting changes requiring the same scrutiny as an API change.

## Why identical-looking prompts miss A prompt cache reuses a stored prefix only when the new request's tokens match the stored ones exactly. Tokens come from bytes, and bytes come from whatever your serializer produced. Nothing in that chain understands that `{"a":1,"b":2}` and `{"b":2,"a":1}` describe the same object. They tokenize differently, so they are different prefixes. Tool and function definitions are the highest-risk content for this because they are usually (a) placed at or near the very front of the reusable region, so any variation invalidates everything behind them, (b) generated programmatically from decorators, registries or reflection rather than written as literal text, and (c) large — a dozen schemas can run to thousands of tokens. ## Where the non-determinism actually comes from **Object key order.** Many serializers emit keys in the order the underlying map yields them. That order can differ between language runtimes, between library versions, after a refactor changes which field is populated first, or — in older or hash-randomised runtimes — between processes of the same build. **Collection ordering.** A tool registry backed by a set, or a plugin loader that walks a directory, produces an order that depends on hashing or filesystem enumeration. Two pods can legitimately register the same eight tools in two different orders. **Conditional inclusion.** Optional fields that are omitted when null in one path and emitted as `null` in another; a `required` array assembled from a set; a description that is only populated when a docstring exists. **Formatting flags.** Indentation, separator spacing, ASCII escaping of non-ASCII characters, trailing newlines. Two services that both "send JSON" can disagree on every one of these. **Value formatting.** Float rendering (`1.0` versus `1`), integer versus string enum values, date formats, and locale-sensitive number formatting. **Semantically irrelevant churn.** Regenerating a schema from a Pydantic-style model after a library upgrade can reorder or rename internal keys such as definition references without any intent to change the contract. ## The diagnostic signature This failure looks different from a volatile field like a timestamp. A timestamp yields a flat 0%. Non-deterministic serialization yields a *partial and unstable* hit rate: fine within one process, poor across a fleet, recovering after a restart, or splitting cleanly along the boundary between two services that both call the model with "the same" tools. Charting distinct prefix hashes per route makes it obvious — you expect one hash and you see a small number, typically proportional to the number of processes or the number of orderings your registry can produce. The confirming step is a programmatic byte diff of two prefixes with different hashes. Visual inspection routinely fails here: reordered keys and a changed float format are almost invisible in a wall of JSON, and a trailing space is fully invisible. ## Canonicalization as the fix Build the reusable region through exactly one function, and give that function these properties: - **Deterministic key order** — sort keys, or emit through an ordered structure you control explicitly. Sorting is usually simplest and is trivially testable. - **Explicit collection order** — sort the tool array by name (or pin an author-chosen order), never rely on registry iteration. - **Fixed formatting** — pin separators, escaping and indentation as constants rather than defaults, so a library upgrade cannot silently change them. - **Normalised values** — one representation per value type; decide once whether an optional absent field is omitted or emitted as null, and apply it everywhere. - **A single owner** — every service that talks to the model imports the same builder. Two copies of "the canonical serializer" drift within a quarter. Then make it verifiable. Hash the canonical string and log the short hash on every call, which gives you the fleet-wide cardinality metric for free. Add a test that constructs the same logical request twice, through two independent code paths if you have them, and asserts the hashes are equal. Add a second test that snapshots the hash for a fixed input, so an upgrade that quietly changes serialization fails a build instead of silently halving your hit rate in production. ## Ordering is a contract, not an implementation detail The deeper lesson a senior interviewer is listening for: once you cache on a prefix, the *serialization* of that prefix has become part of your system's contract. Anything that can reorder or reformat it is now a production-affecting change, even though no behaviour changed and no test failed. Treating tool order as arbitrary — a reasonable stance before caching — becomes a bug afterwards. That is why the guardrail belongs in CI rather than in a code-review checklist: the change that breaks it usually arrives inside a dependency bump that nobody reviewed for prompt-byte stability.

  • How would you distinguish this from a volatile field like a per-request ID as the cause of low hit rate?
    By the shape of the miss. A per-request value gives a flat near-zero hit rate and unbounded prefix-hash cardinality that scales with request volume. Non-deterministic serialization gives a partial, lumpy hit rate with small bounded cardinality — a handful of hashes, often one per process, pod or calling service — and it frequently shifts on restart or deploy. Grouping the hash counts by host or service usually separates the two in one query.
  • What guardrail stops this regressing after a dependency upgrade?
    A test that builds the same logical request twice through independent paths and asserts the canonical prefix hashes match, plus a snapshot test pinning the hash for a fixed input. The second one is what catches library upgrades: a schema generator that reorders internal keys or changes float formatting fails the build instead of silently halving hit rate in production, where it would otherwise show up only as a slow cost drift.
  • Is sorting tool definitions by name always the right canonical order?
    It is the safest default because it is deterministic, stable across runtimes and trivially testable. The caveat is that ordering is not purely cosmetic to the model — position influences selection behaviour, so an author-chosen order may be preferable for quality. Either is fine for caching; what matters is that the order is explicit and pinned rather than inherited from a registry's iteration. Just do not let quality tuning reorder tools per request.

saying these in an interview costs you the question

  • Assuming two JSON objects with identical fields serialize identically
  • Treating map or set iteration order as stable across processes
  • Blaming the provider when hit rate differs between pods
  • Adding retries instead of fixing the serialization
  • Fixing one service's serializer and leaving the others alone

context