skip to content

When a new clause classifier ships, why put the scorer version in the prediction cache key instead of flushing?

level: seniorimportance: should knowfreq 51%

answer

  1. make the rollout the invalidation
  2. scope entries to the scorer that made them
  3. version in the key, not the value
  4. the old namespace survives for rollback
  5. warm the head before the ramp

basics

~20 s

A version in the key makes the rollout itself the invalidation: new-version reads land in an empty namespace and refill, while old entries expire unread. A flush is a separate step that also destroys the entries a rollback would need.

solid answer

~50 s

Put the version in the key and every cached label is scoped to the scorer that produced it - `clause-risk/scorer=<version>/<spanHash>`. Traffic on the new version reads a namespace that is empty and fills it; traffic still on the old one keeps hitting a warm one; and a rollback is a routing change rather than a refill, because the old namespace was never destroyed. A flush does none of that. It has to be coordinated with the deploy as a separate action, it removes the old version's entries too, so a rollback lands on a cold cache exactly when the system is already unhealthy, and it hands the scoring fleet its entire miss load at once instead of per key as the ramp grows. Old entries need no cleanup: once nothing reads them they expire under their own TTL.

code

json · 12 lines
json
{
  "key": "clause-risk/scorer=2026-09-11.a3f/sha256:9f2c4be1d0a7...e41",
  "value": {
    "labels": [
      { "name": "limitation-of-liability", "score": 0.94 },
      { "name": "indemnity", "score": 0.11 }
    ],
    "scoredAt": "2026-09-17T08:59:11Z",
    "featureAsOf": "2026-09-17T08:00:00Z",
    "expiresAt": "2026-09-17T13:59:11Z"
  }
}

go deeper

for a junior

Know that a cached prediction belongs to the model that produced it, so deploying a new model does not by itself change what a cache hit returns.

for a middle

Explain namespacing by scorer version: new traffic reads an empty namespace and fills it, old entries expire unread, and nothing has to be deleted by hand.

for a senior

Argue the rollback and the load case - a flush destroys the entries you need most during a bad release and hands the scoring fleet its whole miss load in one instant.

for a principal

Own what the version token identifies: the entire scoring function including preprocessing, so that a key means one stable input-to-output contract across releases, ramps and rollbacks.

## What a rollout does to a cache that knows nothing about versions The scoring service is replaced; the cache is not. Requests that hit keep returning labels the retired classifier produced, and they keep doing it until each entry's TTL runs out. Teams meet the consequences in this order: 1. **The new classifier only sees the misses.** At a 55% hit rate, 45% of traffic is scored by the new model and the rest is answered by the old one's stored output. 2. **The change appears to arrive gradually.** Instead of switching at deploy time, behaviour drifts toward the new model over roughly one TTL as entries expire - which makes any reading taken over that traffic understate the change. 3. **A targeted fix does not reach the people who reported it.** The clause that was mislabelled is cached under exactly the key those users hit, so the fix ships and the complaint stands. ## Version in the key, not in the value Lay the key out as a namespace: `clause-risk/scorer=2026-09-11.a3f/sha256:9f2c...`. Four properties follow directly from that shape: - Reads from the new version address an **empty namespace** and fill it as traffic arrives, per key, in proportion to the ramp. - Reads still on the old version hit a **warm namespace**, so a ramp does not degrade the traffic that has not moved yet. - **Nothing is deleted at deploy time.** Old entries stop being read and expire under their own TTL. - **Rollback is a routing change.** Traffic returns to a namespace that is still populated, so recovery costs no scoring at all. Keeping the version only in the entry's value is not unworkable, but it is strictly worse: every request must read the entry and then discard it on a mismatch, one key can hold only one version's answer, so during a ramp the two versions overwrite each other's entries, and a rollback starts from an empty cache. ## What the version must actually identify The **whole scoring function**, not the release ticket: the model artifact, the preprocessing, the span segmentation and the label set it was trained with. If segmentation changes and the version does not, the same span hash now denotes a different input, and the cache confidently returns an answer computed for text that was never sent. Deriving the version from a content hash of the served bundle is safer than a tag typed by a human, because it cannot be forgotten. ## The cold namespace, and warming it A fresh namespace starts empty, which is the one real cost of this design: at the moment traffic moves, every request is a fill. Because the frequency falloff is steep, a small warming job removes most of that cost - replay the few thousand most requested spans from yesterday's trace through the new version, rate-limited and off-peak, before the ramp begins. Two rules for the warming job: - **Score them; never copy them.** Copying old entries into the new namespace publishes the retired model's output under the new model's key, which is a silent wrong answer of the worst kind. - **Warm the head, not the tail.** The tail costs nearly a full day of inference to prepare keys that may not be requested again during the ramp. ## Two namespaces and what they cost During a ramp both namespaces exist, but neither is a full copy: each holds only the keys its own traffic has touched, and since the head dominates the distribution, the extra memory is roughly the head of the key space rather than a second complete cache. Once the old version is retired its entries expire on their own, or the namespace is dropped in one operation. ## When a flush is still the right call The argument above is about the ordinary case, where the outgoing model was merely older. It inverts when the outgoing model was actively harmful - a bug that mislabels a whole clause family, output that must not be served for another minute. Then you do want those entries gone immediately, and the namespace layout is what makes that cheap: dropping one namespace is a single operation rather than a scan for matching keys. Versioned keys do not remove the ability to flush; they remove the *need* to flush on every ordinary release.

  • What exactly should the version token in the key identify?
    The whole scoring function rather than the release: the model artifact plus the preprocessing, the span segmentation and the label set it was trained with. If segmentation changes while the token does not, the same span hash denotes a different input and the cache serves an answer computed for text nobody sent. Deriving the token from a content hash of the served bundle makes it impossible to forget.
  • Do both namespaces have to be sized for full traffic during a ramp?
    No - only the overlap needs room. Each namespace holds the keys its own share of traffic has touched, and the head of the distribution dominates, so the extra memory is roughly the head rather than a second full cache. When the old version is retired, its entries expire under their own TTL or the namespace is dropped outright.
  • How would you confirm after the deploy that cached traffic really moved to the new scorer?
    Count hits and fills per namespace. The old version's hit count should fall to zero as its traffic moves, the new one's fills should rise and then level off as it warms, and the two should never be served concurrently to the same request path. A hit count that keeps ticking on the retired namespace means some caller is still constructing the old key.

saying these in an interview costs you the question

  • Ships a new scorer and expects cached traffic to pick it up.
  • Flushes the entire cache at deploy time and calls it invalidation.
  • Stores the version in the value while still keeping one entry per span.
  • Versions the deploy rather than the scoring function, so preprocessing slips through.
  • Warms a new namespace by copying the old version's entries under the new key.
  • Assumes the old namespace must be deleted the moment the ramp completes.