Why would a photo library precompute object tags nightly but compute person tags on the request?
answer
- split by knowability, not by cost
- facts about the item versus the viewer
- the pass cannot know the request
- the heavy half can still be precomputed
- bound the rescore to the screenful
basics
~20 sBecause the split follows what the request knows that the scheduled pass could not. Object tags depend only on the pixels, so they can be produced once; person tags depend on how this owner has named their face clusters, which changes after the pass ran.
solid answer
~50 sA hybrid is not a compromise between two placements, it is a split along one line: the part of the prediction whose inputs were fully known when the pass ran goes into the precomputed store, and the part that depends on something the request supplies is produced on the read. Objects and places are a property of the pixels and never change, so precomputing them is free of staleness. A person tag resolves a detected face against the owner's own clusters and names, and that mapping changes whenever the owner merges or renames a cluster — a precomputed row would be wrong the moment they did. The rescore must be bounded by what is on screen, typically a screenful rather than a 5,000-photo album, or you have quietly moved the whole album into the read path.
go deeper
Know that some parts of a prediction depend only on the item and some depend on who is looking, and that only the first kind can be worked out in advance.
State the split rule — precompute what the pass could know, resolve on the request what only the request knows — and give an example where the expensive half is the precomputed one.
Bound the request-time half explicitly by the screenful and by the stored artefact it consumes, and define the version contract between the two halves so an incompatible pair is caught rather than composed.
Weigh the freshness a hybrid buys against running two placements with two failure modes and two version lines. If nothing in the prediction genuinely depends on the request, the simpler single placement is the better system.
## Why a hybrid exists at all Splitting a prediction across two placements looks like indecision and usually is not. It exists because a single prediction is often a composition of parts whose inputs become knowable at different times. Precompute can only capture the parts whose inputs were complete when the pass ran. Anything that depends on the request — who is asking, what they have configured since, what they did five minutes ago — was not available to the pass and cannot be in the stored row at any cadence. ## The split rule The rule is one sentence: **precompute what the pass could know; compute on the request what only the request knows.** Two useful consequences follow. - It is not a cost split. "Put the expensive half in the batch job" sounds sensible and produces wrong answers, because expense says nothing about whether the stored value will still be correct when it is read. - It is a staleness split. A precomputed part whose inputs cannot change is never stale except against the scorer version. A part whose inputs change after the pass is stale the moment they do, and no cadence short of the read fixes it. ## The photo library worked through | part of the tag | what it depends on | placement | |---|---|---| | objects and scene | the photo's pixels, fixed at upload | precomputed, nightly pass or on the upload event | | place | the pixels plus capture metadata, fixed at upload | precomputed | | a detected face region and its embedding | the pixels, fixed at upload | precomputed, the expensive part | | which person that face is | the owner's current face clusters and names | computed on the request | The heavy work — detecting faces and producing an embedding per face — is done once, in the batch pass, and stored. The cheap work — matching that stored embedding against the owner's current clusters and resolving a name — runs on the read, because the owner may have merged two clusters or corrected a name since the pass. Note that the expensive part landed in precompute and the cheap part on the request, which is exactly the opposite of a cost-driven split and exactly what the staleness rule predicts. ## Bounding the request-time part The part that runs on the read has to be bounded by what the viewer actually sees: 1. **Bound by the screenful, not the album.** An album of 5,000 photos renders perhaps sixty at a time. Resolving names for sixty stored embeddings is a small, predictable amount of work; doing it for 5,000 is the whole album moved onto the read path. 2. **Bound by the stored artefact.** The request-time step consumes the precomputed embedding rather than the photo, so it never re-runs the heavy detector under a user-facing deadline. 3. **Bound by the owner's own data.** The cluster set belongs to one library, which keeps the request-time comparison small no matter how large the corpus is. If any of those bounds is missing, the hybrid has stopped being a hybrid and has become live scoring with extra steps. ## When a hybrid is the wrong answer - **Nothing in the prediction depends on the request.** Then it is a pure precompute problem, and adding a request-time step buys complexity with no freshness. - **The request-time part cannot be bounded.** If resolving the viewer-dependent piece requires touching the whole album or the whole corpus, the split has not actually removed the work from the hot path. - **The two parts must be consistent with each other.** If the precomputed part and the request-time part are produced by versions that disagree — an embedding from one detector, a cluster set built from another — the composition is wrong in a way neither half reveals. The stored artefact's version has to be part of the contract between them. - **The operational cost exceeds the freshness gained.** Two placements mean two failure modes, two things to size and two versions to keep compatible. That is worth paying when the request-time input genuinely changes; it is not worth paying to shave milliseconds. ## The sentence to say in a round "I precompute everything that is a fact about the photo, and I resolve on the request everything that is a fact about the viewer — bounded to what is on screen, and reading the precomputed artefact rather than the photo." That single line contains the split rule, the bound and the contract between the two halves.
- Why is splitting by cost per item the wrong instinct?Because cost says nothing about whether a stored value stays correct. In a photo library the expensive step — face detection and embedding — is perfectly precomputable since the pixels never change, while the cheap step — resolving which person a face is — depends on naming the owner changed yesterday. Cost-driven splitting lands both on the wrong side.
- What contract has to hold between the precomputed half and the request-time half?A version contract. The request-time step consumes a stored artefact, so it must know which producer made it and refuse or recompute when the versions are incompatible. Without that, an embedding from an old detector is silently compared against clusters built from a new one, and the composed answer is wrong in a way neither half can detect.
saying these in an interview costs you the question
- Splits the work by cost per item rather than by what the request knows
- Rescores the whole album instead of the photos currently on screen
- Says the precomputed half must be rebuilt whenever the viewer changes anything
- Claims a hybrid is always better than either pure placement
- Lets the request-time step re-read the photo instead of the stored artefact