In a photo library where only one stored photo in five is ever opened, why precompute tags for all 400 million?
answer
- count distinct photos, not reads
- uploads per day set the recurring bill
- two read shapes, only one is live-able
- a ranging read has no candidate list
- capability first, then cost per call
basics
~20 sBecause tag search reads across the whole library without naming any photo, and only rows produced in advance can answer it. Four-fifths of the precomputed work is never read, and that waste is accepted, not denied.
solid answer
~40 sCount the calls first: at 20 million uploads a day, tagging on the upload event costs 20 million scorer calls daily, while tagging on first view costs only the fifth that gets opened — about 4 million. On call count alone, live scoring wins by five to one. It loses anyway, because the product has a second read shape: a tag search names no photo ids and must consider all 400 million rows, and a scorer cannot be called for keys the request never supplied. So the argument is not `cheapest placement` but `which placements can serve every read path` — and only after that, which of the survivors is cheaper. The 80% of precomputed rows nobody reads is a known, priced cost of making search possible.
code
pseudocode · 19 linesCORPUS = 400_000_000 # photos already stored
UPLOADS_DAY = 20_000_000 # photos arriving per day
OPEN_RATE = 0.20 # fraction ever opened by anyone
function calls_per_day(placement):
if placement is NIGHTLY_FULL_PASS:
return CORPUS # 400M, rescored every night
if placement is ON_UPLOAD_EVENT:
return UPLOADS_DAY # 20M, once per photo ever
if placement is ON_FIRST_VIEW:
return UPLOADS_DAY * OPEN_RATE # 4M, only opened photos
function serves_tag_search(placement):
# a search supplies no photo ids, so nothing can be scored for it
if placement is ON_FIRST_VIEW:
return false
# the two precompute triggers qualify only once the stored
# backlog has had one full pass over CORPUS
return truego deeper
Know that precomputing means paying for photos nobody will look at, and that the fraction ever opened is the number that says how much. Be able to multiply an upload rate by that fraction.
Work the arithmetic out loud in scorer calls per day per placement, and separate the recurring bill (uploads) from the one-off (the stored backlog). Name the read that cannot be served live.
Order the decision correctly: capability first, cost second. Say which reads range over unnamed keys, eliminate the placements that cannot answer them, and then state the accepted waste as a number the team signs off on.
Ask whether search has to depend on the model at all. Coupling a corpus-wide capability to an expensive scorer means every future model change is a corpus-sized job, and a cheaper coverage signal may be worth more than the tags it replaces.
## The waste is real, and it is not the argument In a library of 400 million stored photos where roughly one in five is ever opened, precomputing a tag row for every photo means four rows in five are written, stored and eventually rewritten without a single read. That is genuinely wasteful, and a candidate who pretends otherwise is not being honest about the design. The reason to do it anyway is that one of the product's read paths cannot be served any other way. ## Count the calls honestly Use the steady-state upload rate, not the corpus size, for the recurring bill: - **Nightly full pass:** 400,000,000 scorer calls per day. Almost all of them rescore a photo whose pixels have not changed. - **Tagging on the upload event:** 20,000,000 calls per day — one per arriving photo, once, forever. - **Tagging on first view:** 20,000,000 x 0.20 = 4,000,000 calls per day in steady state, because only the opened fifth is ever scored. The common error here is counting *reads* instead of *distinct photos*. Thirty million album-open reads a day do not mean thirty million scoring calls under a first-view placement: a photo scored once has a row, and subsequent opens are lookups. ## What the call count leaves out | cost | batch / event precompute | live on first view | |---|---|---| | scorer calls per day | 20M (event) | 4M | | rows stored | 400M, including unread ones | only what was opened | | work when a new scorer ships | rescore the corpus | nothing, next read uses it | | scoring inside a user deadline | no | yes, on the first open | | can serve a search over unnamed keys | yes | no | The last row is the one that decides. The others trade against each other; that one is a capability, not a price. ## The read pattern decides before the cost curve Sort the product's reads into two classes: 1. **Reads that name their keys.** Opening an album hands the service a list of photo ids. Every placement can serve this, because a live scorer knows exactly what to score. 2. **Reads that range over keys the request never names.** "Show me every photo with a beach in it" is a query over the whole library. A live scorer has nothing to score against — the candidate set *is* the corpus. This read can only be answered from rows that already exist. Once the product has a tag search box, precompute stops being an optimisation and becomes a requirement for the corpus, and the remaining questions are about the trigger and the backlog. ## The backlog is a separate line item Tagging on the upload event covers arrivals only. The 400 million photos already stored when the placement is chosen are not covered by any future event, so event-driven tagging needs exactly one historical backfill pass before search works over the whole library. That pass is a batch job — which is why real systems usually run both triggers: a one-off (or per-model-version) pass over the corpus, plus a standing consumer on the upload stream. ## When the same arithmetic flips the other way The conclusion is not "always precompute". Change one number and the answer changes: - **No ranging read.** If tags are only ever shown on a photo the request already names, the capability argument disappears and 4 million calls beats 20 million. - **A very expensive scorer.** If the per-call cost is high enough, paying it for 80% of rows nobody reads can exceed the value of search, and the honest design tags only photos in albums the owner has actually browsed. - **Inputs that change.** If the prediction depends on something that moves after upload, a stored row starts rotting and precompute buys a stale answer rather than a cheap one. - **A read that is not user-facing.** If search is served from a periodically rebuilt index rather than the live path, the freshness requirement on the precomputed rows loosens considerably. ## The rule to state out loud 1. List the read paths and mark which ones range over keys the request does not name. 2. Eliminate placements that cannot serve one of those paths. 3. Only now compute scorer calls per day for the survivors, using the upload rate and the fraction ever opened. 4. Add the non-recurring costs — the historical backfill, and one corpus-sized rescore per model version. 5. State the accepted waste as a number, so nobody rediscovers it later as a surprise.
- First-view tagging is five times cheaper in calls. What would make you choose it despite search?Serving search from a separate, cheaper signal — filename, capture metadata, an owner's manual albums — and reserving the model's tags for photos actually viewed. That is a different product promise, and it is the honest trade: you are not making precompute cheaper, you are removing search's dependency on the scorer.
- Does storing 400 million tag rows change the arithmetic much?Usually far less than the scoring does. A tag row is a short list of identifiers and scores, so the corpus is a modest amount of structured data next to the photos themselves. Storage is worth stating as a number, but the placement decision is rarely settled by it.
- How would you reduce the 80% waste without giving up search?Score everything once with a cheap tagger to guarantee coverage, and reserve the expensive scorer for photos that enter a read path — opened, shared or matched by a search. Coverage stays complete, and the costly work follows actual demand rather than the corpus.
saying these in an interview costs you the question
- Picks the placement with the fewest scorer calls per day and stops
- Counts one scoring call per album-open read instead of per distinct photo
- Assumes a precomputed row is free once it has been written
- Says an 80% unread rate rules precompute out by itself
- Thinks a search over the library can call the scorer per query
- Forgets the existing corpus needs a backfill under event-driven tagging