A nightly tagging job leaves today's photo uploads untagged when an album opens — what does that force into the read path?
answer
- precompute always leaves uncovered keys
- reads skew recent, coverage does not
- miss rate is a read statistic
- size the live tier from miss traffic
- concurrency is arrival rate times service time
basics
~20 sIt forces a live scorer onto the read path for the uncovered photos, which makes the design a hybrid. That live tier has to be sized for the miss rate, and album opens skew hard toward recent photos.
solid answer
~50 sChoosing precompute does not remove request-time scoring; it decides how much of it you do. Any photo the pass has not reached is a coverage miss, and the album still has to render, so the read path looks up the tag store, collects the ids with no row, scores them live and writes them back so the next open is a lookup. Size that tier from the miss traffic, not from zero: with 30 million album-open reads a day and roughly 40% of them landing on photos uploaded since the last pass, the live path carries about 12 million calls a day — near 140 per second on average, and perhaps 420 at peak. By Little's Law at a 200 ms service time that is around 28 scorings in flight on average and 84 at peak. Nobody provisions that by accident.
code
pseudocode · 14 linesfunction on_album_open(photo_ids):
rows = tag_store.lookup_many(photo_ids) # written by the last pass
missing = photo_ids for which no row came back
if missing is empty:
return rows # the intended fast path
# every miss is a scoring call inside the album-open deadline
scored = live_scorer.score(missing)
for each photo_id, tags in scored:
tag_store.write(photo_id, tags, scorer_version)
emit_metric("tag_miss_ratio", size(missing) / size(photo_ids))
return rows merged with scoredgo deeper
Understand that a scheduled pass leaves recently uploaded photos without tags, and that something has to happen when an album containing them opens.
Describe the read path concretely: look up the store, collect the ids with no row, score those live, write them back. Explain why the write-back stops the same photo being scored on every visit.
Size the live tier from the miss traffic rather than the uncovered fraction, show the Little's Law concurrency at the peak, and treat the cadence as the knob that trades batch work against that fleet.
Decide whether the product can carry a heavy scorer on a user-facing read at all. If the peak miss load will not fit an album-open deadline, the scheduled cadence is not a viable placement and the arithmetic says so before anyone builds it.
## A precompute placement is a hybrid whether you designed it or not The sentence "we precompute the tags" implies the read path never scores. It does, for every key the pass has not covered yet, and in a photo library that set is not marginal. A nightly pass covers photos as of the moment it read them; everything uploaded since is uncovered, and the viewer opening an album does not care why. Either the album renders without tags, which is a product decision rather than an architecture, or the read path scores the uncovered ids. The moment you choose the second, you own a live scoring tier and have to size it. ## The misses are not a random sample of the corpus The tempting arithmetic is: the pass covers everything except one day's uploads, 20 million out of 400 million, so 5% of reads miss. That is wrong, and it is wrong in the expensive direction. Album opens are heavily recency-skewed — people look at the trip they took last weekend far more than at photos from four years ago. If 40% of album-open reads land on photos uploaded since the last pass, the miss rate on the read path is **40%, not 5%**, even though uncovered photos are 5% of the corpus. The distribution of reads over keys, not the fraction of keys covered, is what sets the load. This is the same shape as any heavy-tailed access pattern: a small slice of keys takes a large share of reads. Assume it and measure it; do not assume reads are uniform over the library. ## Sizing the live tier Work it through end to end: - 30,000,000 album-open photo reads per day = 30,000,000 / 86,400 ~= **347 reads per second** average. - 40% of them hit an uncovered photo = 12,000,000 per day = **~139 live scoring calls per second** average. - A diurnal peak of roughly 3x average gives **~417 calls per second** at peak. - **Little's Law** — concurrency = arrival rate x service time — at a 200 ms service time gives 139 x 0.2 ~= **28 scorings in flight** on average and 417 x 0.2 ~= **84 in flight** at peak. That is the fleet the "precompute" design actually requires. A team that sized for zero because the tags are precomputed discovers the number during the first busy evening. ## Write back, or pay for the same photo repeatedly When the live path scores an uncovered photo, it should write the row into the same tag store the pass writes to. Two things follow: 1. **The miss is paid once per photo, not once per open.** Without the write-back, a popular recent album re-scores the same photos on every visit, and the live tier's load tracks reads rather than distinct photos. 2. **The pass has less to do.** Rows already produced at read time do not need producing again, so the boundary between the two paths is a fill-order question rather than a duplication. The row has to carry the scorer version that produced it, so that a later pass can tell a current row from one produced by a superseded model. ## What this does to the placement decision Once you accept that precompute plus an on-demand fill is one design rather than two, several things follow: - **Cadence is a load knob.** Running the pass more often shrinks the uncovered window and therefore the live tier, at the cost of more batch work. The two costs are directly comparable. - **An arrival-triggered consumer collapses the miss rate.** Tagging on the upload event covers a photo seconds after it lands, so the read path misses only on the backlog and on genuine gaps, not on everything uploaded today. - **The live tier's ceiling is a design constraint.** If the scorer is too heavy to run inside an album-open deadline at 400 calls per second, then a nightly cadence is not a viable placement for this product, and that conclusion arrives from the arithmetic rather than from taste. - **A miss is not an error.** An uncovered photo is the expected steady state of a scheduled placement; treating it as a failure hides the real signal, which is the miss *rate* trending against the cadence you chose. ## The number to watch afterwards Instrument the read path with the fraction of requested ids that had no row. That single ratio tells you whether the pass cadence still matches the upload rate, catches a stalled pass long before anyone complains about missing tags, and is the input to every later decision about whether to shorten the cadence or move to arrival-triggered tagging.
- Why is the miss rate 40% when uncovered photos are only 5% of the library?Because reads are not uniform over keys. Album opens concentrate on recent uploads, so the uncovered slice of the corpus absorbs a share of reads far larger than its share of rows. Coverage is a property of the corpus; miss rate is a property of the read distribution, and only the second one sizes the fleet.
- What happens to the live tier if you halve the pass cadence?The uncovered window halves, so the share of reads landing on uncovered photos falls and the live tier shrinks roughly in proportion. You pay for that with a second full pass per day, so the two are directly comparable in cost and the cadence becomes a tuning knob rather than a habit.
- Should the live path write its result into the same tag store the pass writes?Yes, stamped with the scorer version that produced it. Without the write-back the same recent photos are rescored on every visit and the live tier's load tracks reads rather than distinct photos. With it, each uncovered photo costs one live call in its lifetime.
saying these in an interview costs you the question
- Assumes a precompute placement means the read path never scores
- Sizes the live scorer for zero traffic because tags are precomputed
- Sets the miss rate to the uncovered fraction of the stored library
- Treats an uncovered photo as a scorer failure rather than expected
- Calls write-back optional bookkeeping instead of what caps repeat misses
- Divides daily traffic by a day twice and reports a tiny miss rate