A legal-document classifier caches finished labels; why does keying on upload id nearly never hit while a content hash does?
answer
- the key decides everything
- what actually repeats in the traffic
- uploads are unique, text is not
- hash the normalised clause span
- normalise only what the scorer ignores
basics
~20 sAn upload id is unique to one upload, so every lookup misses. A content hash collapses the identical templates and boilerplate clauses that recur across customers onto one key, and repeats are the only thing a prediction cache can serve.
solid answer
~40 sA prediction cache pays only where the same input recurs, so the key has to name what recurs. In a review workflow the document is new every time: an `uploadId` key gives a hit rate near zero however large the cache is, while still costing a lookup and a write on every request. The text is not new. Standard contract types and boilerplate clauses repeat across customers and across revisions of the same template, so hashing the normalised clause span makes those identical spans collide onto one key and turns scoring into a lookup for all of them. The key must also pin everything else the prediction depends on that the hashed text does not carry - the scorer version, the label set, and any per-customer input.
go deeper
Recall that a cache only helps when the same lookup key comes back. A per-upload identifier is new on every request, so nothing repeats and the cache returns nothing while still costing a read.
Explain how the unit of caching is chosen: hashing normalised clause spans makes boilerplate collide across customers, and the normalisation has to match what the scoring path already does to its input.
Show that you would check the key against the real inputs - scorer version, label set, anything customer-specific - before letting a globally shared content hash answer one tenant with another tenant's prediction.
Frame it as a data-sharing decision as much as a performance one: a content-addressed prediction cache is a cross-tenant surface, and the guarantee that a prediction is a pure function of the text is what makes that sharing safe.
## What this cache is doing A **prediction cache** stores the output of a scoring call - here the contract type, clause categories and risk flags a classifier produced for a piece of legal text - so that a later request for *the same input* is answered by a read from a low-latency key-value store instead of a forward pass. It is not a generic response cache and not the feature cache behind the model: it holds finished model output, addressed by whatever identifies the input. That makes the cache's value one number, the **hit rate**, and the hit rate is decided almost entirely by the key. Capacity and eviction move it at the margin; identity decides whether repeats exist at all. ## What actually recurs in this workload Traffic in a legal-document review workflow has a very particular shape: - **Uploads are unique.** Every submission gets a fresh identifier, and most customers upload a given document once. - **Text is not unique.** Contracts are assembled from templates. A limitation-of-liability clause, a governing-law clause or a standard confidentiality body appears verbatim in thousands of documents and across many customers. - **Re-review is common.** The same text is scored again when a reviewer reopens a file, when a label is added to the taxonomy, or when a sweep re-runs a customer's back catalogue. A key built on the upload identifier records only the first fact. Every request carries a new id, every lookup misses, and the cache returns nothing while still costing a round trip on the read and another on the write. A key built on a **content hash** records the second and third: identical text collapses to one entry no matter who sent it or when. | Key | What collides | Hit rate | Watch out for | |---|---|---|---| | `uploadId` | nothing | about 0% | pure overhead; the entry is never read again | | `hash(document)` | byte-identical documents | low | one changed date makes a new key | | `hash(clause span)` | boilerplate across customers | high | the path must segment before scoring | | `(customerId, templateId)` | documents from one template | moderate | wrong when filled-in fields drive the label | The third row is why this subject exists. Scoring **per clause span** moves the unit of reuse down to the thing that actually repeats: a forty-page contract can be entirely unique as a byte string while thirty-four of its fifty clauses are boilerplate the classifier has already scored a thousand times. ## Normalisation is part of the key Two spans differing by a trailing space should not be two keys, so the text is normalised before hashing - whitespace collapsed, line breaks folded, quote characters and soft hyphens canonicalised. The rule that keeps this honest is: **normalise only what the scorer itself ignores.** If the key-building step lowercases but the model is case-sensitive, the cache will answer one span's request with another span's label. That is a silent wrong answer, not an error. Every normalisation applied to the key must be one the scoring path already applies to its own input, or the entry stops being a faithful memo of the scoring call. ## What else belongs in the key The key names the *whole* input to the scoring function, not just the text: 1. **The scorer version.** The same clause scored by a new classifier is a different answer. An entry keyed without it is served after a rollout, and the rollout quietly does nothing for cached traffic. 2. **The task or label set.** One encoder feeding a contract-type head and a risk-flag head produces two different outputs for one span. 3. **Any input outside the hashed text.** If the risk flag also depends on the customer's account tier, the tier belongs in the key or the cache will serve one customer's flag to another. Point three is also the tenancy question. A content hash is global by construction, which is exactly what you want when the prediction is a pure function of the text: shared boilerplate is scored once for everybody. It becomes a correctness bug, and potentially an exposure of one customer's circumstances to another, the moment the prediction depends on anything customer-specific. If it does, scope the key or do not cache that output. ## Collisions and the honest failure mode A hash-keyed cache is wrong if two different inputs hash alike. Use a wide cryptographic digest - `sha256` is the usual choice - and resist truncating it to save memory. A short digest over hundreds of millions of spans makes a collision a realistic event, and its symptom is one clause wearing another clause's label with nothing anywhere in the system reporting a problem. Store the digest at full width; if memory is the constraint, store fewer entries rather than shorter keys.
- A reviewer edits one clause and resubmits the document - what happens in the cache?Only the edited span misses. Every other span's hash is unchanged, so the resubmission is answered almost entirely from the cache and the scorer sees a single span. That is the practical argument for span-level keys over a whole-document hash, where one changed character makes the entire document a new key and re-scores fifty clauses to learn one of them moved.
- Should the same key be used for a customer-specific prediction and a shared one?No - keep them in separate namespaces. A prediction that is a pure function of the text can be shared globally and is where the hit rate comes from. One that also reads an account tier or a customer's history must carry that input in its key, which fragments the key space and lowers the hit rate; mixing the two under one key shape is how a shared entry ends up answering for the wrong tenant.
saying these in an interview costs you the question
- Thinks a bigger cache fixes a hit rate near zero.
- Keys on the upload id because it is already the primary key.
- Hashes raw bytes and expects boilerplate in different files to collide.
- Normalises aggressively without checking the scorer does the same.
- Leaves the scorer version out of the key and flushes by hand at release.
- Truncates the digest to save memory and calls a collision impossible.