Why can cached clause embeddings survive a retrain that invalidates every cached label?
answer
- two points on the same path
- different keys, different versions
- the encoder is the stable half
- an embedding hit shortens the miss
- vectors cost kilobytes, labels cost bytes
basics
~20 sBecause the two caches name different versions in their keys. An embedding depends only on the encoder, a finished label on the classifier above it, so refitting the head invalidates the labels and leaves every embedding entry valid.
solid answer
~50 sThere are two places to cache on a scoring path. The embedding cache stores the encoder's output vector for a normalised clause span, keyed by span hash plus encoder version. The prediction cache stores the finished labels, keyed by span hash plus the version of the whole scoring function. Most retrains in a review workflow change the classifier above the embeddings - a new label added to the taxonomy, a head refitted on recent reviewer verdicts - while the encoder stays fixed. That bumps the prediction namespace and leaves the embedding namespace untouched, so the cold label cache refills cheaply: the expensive part of the forward pass is already stored. The trade is size, because a vector is far larger than a handful of labels, and the day the encoder is replaced every entry is invalid at once.
go deeper
Know that a scoring path has stages, and that a cache can store either the finished answer or an intermediate result produced on the way to it.
Explain the two keys: a vector depends on the encoder version alone, while a label depends on every component of the scoring function above the encoder.
Weigh the memory and the fetch: a stored vector is orders of magnitude larger than a label, so this only pays back where the encoder is genuinely stable and expensive.
Decide where the reuse boundary sits across teams - a shared vector namespace lets several downstream heads reuse one costly computation, at the price of making the encoder version a contract they all depend on.
## Two places to cache on one path A clause-scoring request runs through stages: segment the document, encode each span into a vector, apply a classifier head to that vector, post-process into labels and risk flags. A cache can sit at the end of that path or in the middle of it, and the two choices behave differently under change. | What is stored | Key pins | Invalidated by | Entry size | Reuse | |---|---|---|---|---| | finished labels | span hash + scoring-function version | any change to encoder, head, label set or preprocessing | tens of bytes | one consumer | | the span's vector | span hash + encoder version | a change to the encoder only | kilobytes | every head reading that encoder | The difference in the second column is the whole answer. A label is the output of everything; a vector is the output of one stage. The narrower the dependency a key pins, the more changes the entry survives. ## Why the encoder is usually the stable half In a document-review workflow, the things that change often are the ones close to the business: a new clause category is added to the taxonomy, a head is refitted on the last quarter's reviewer verdicts, a risk threshold moves. The text encoder is expensive to train, is shared by several heads, and is replaced rarely. So the common retrain bumps the scoring-function version, empties the prediction namespace and leaves the embedding namespace exactly as it was. That is a condition, not a law. Where the encoder is fine-tuned on the same cadence as the head, the embedding cache is invalidated just as often and buys nothing but memory pressure. Check which half of your stack actually moves before designing around this. ## What an embedding hit actually buys An embedding hit does not remove the scorer from the request path; it shortens it. A request whose label is a miss but whose vector is a hit skips the encoder pass and still runs the head, which is the cheap end of the computation. So the payoff is proportional to how much of the forward pass the encoder represents - substantial for a deep text encoder over a long clause, negligible where the model is small. The second payoff is across consumers. One namespace of span vectors can serve a contract-type head, a clause-category head and a risk head at once, so an expensive encode is paid for once and reused three times. The price is that the encoder version becomes a contract between those consumers: whoever replaces it invalidates everyone's entries simultaneously. ## The arithmetic of storing vectors Size is what usually decides this. A vector of 768 dimensions stored as 32-bit floats is about 3 KB. Holding one for each of 600,000 distinct spans is roughly 1.84 GB, against perhaps 60 MB for the same spans' labels. Two consequences follow: 1. **The same memory holds far fewer embeddings than predictions**, so the embedding cache is normally scoped to the head of the key distribution while the prediction cache can cover more of the tail. 2. **Fetching an entry is no longer free.** Moving 3 KB across the network costs real time, and for a short span on an idle accelerator, re-encoding can genuinely be cheaper than the lookup. Measure the two before assuming the cache wins. Storing the vectors at reduced precision halves the footprint and changes the stored value, so it has to be treated as a change to what is cached rather than a storage detail. ## A different question from putting embeddings in the feature store Whether an embedding belongs in the online store as a model *input* - how it is refreshed, who owns its definition, what a stale vector means as a feature - is a separate subject. This is narrower: a memo of a deterministic function, so that the same span is not encoded twice. If the vector is a feature with a life of its own, the version-and-TTL reasoning above is not the mechanism governing it. ## When not to bother Skip the embedding cache when the encoder is retrained as often as the head, when the encode is cheap relative to a multi-kilobyte fetch, or when the key space is so flat that few spans are ever encoded twice. Skip it also when only one head consumes it and the finished-label cache already covers that traffic - in that case you are paying vector-sized memory for a benefit the smaller cache is already delivering.
- What has to be in an embedding cache key besides the span hash?The encoder version and the preprocessing that produced its input - the normalisation and the segmentation rule. Nothing else, and deliberately not the head, the label set or the decision thresholds, because pinning those is what would tie the entry's life to changes that do not affect the vector at all.
- If a request's label misses but its embedding hits, what work is skipped?The encoder pass over that span, and nothing else. The head still runs over the stored vector and the post-processing still runs over its output. That is why the saving is proportional to the encoder's share of the forward pass: it is large for a deep text encoder over a long clause and close to nothing for a small model.
saying these in an interview costs you the question
- Assumes any retrain invalidates the embedding cache as well.
- Caches embeddings without the encoder version in the key.
- Expects an embedding hit to skip the classifier head too.
- Sizes an embedding cache as if entries were the size of a label.
- Caches embeddings when the encoder is refitted as often as the head.