skip to content

In a photo library that auto-tags every upload, how do nightly batch tagging, tagging on the upload event, and tagging on first view differ?

level: juniorimportance: must knowfreq 72%

answer

  1. where in time the scorer runs
  2. clock, arrival, or read
  3. batch and event are both precompute
  4. coverage against work nobody reads
  5. a ranging read cannot be scored live

basics

~10 s

They differ in when the scorer runs. A nightly job tags the stored library on a schedule, event tagging scores each photo as it arrives, and first-view tagging scores only photos somebody actually opens.

solid answer

~50 s

All three produce the same tag; they differ in when the scorer runs, and therefore in coverage, freshness and wasted work. A nightly pass walks the stored library and writes a tag row per photo, so a photo uploaded at noon carries no tags until the job reaches it — the age of a tag is bounded by the cadence, not by the upload. Event-driven tagging runs one scoring call per upload as it lands, so every new photo is covered within seconds, but every photo is paid for whether or not anyone opens it. First-view scoring pays only for opened photos and always uses the deployed scorer, but it puts scoring inside a user-facing read and leaves the rest of the library with no tags at all — which breaks any read that ranges over photos the request never named, such as tag search.

go deeper

for a junior

Be able to name the three placements and say when a tag first exists under each. The compact version is: a clock triggers it, an arrival triggers it, or the read triggers it.

for a middle

Explain why batch and event-driven are the same placement family with different triggers, and why only request-time scoring sits inside a user-facing deadline. Put numbers on the scorer calls each one makes per day.

for a senior

Drive the choice from the product's read shapes rather than from cost alone: a read that ranges over keys it never names settles the argument before any cost curve does. State what your choice leaves broken.

for a principal

The lasting cost is not the calls, it is the commitment: a precomputed store makes every future model change a corpus-sized job, and that bill lands on whoever owns the platform, not on the team that chose the placement.

## The question behind the question A prediction has to exist somewhere before it can be shown. **Placement** is the design decision about *where in time* the scorer runs relative to the read that needs its output. In a consumer photo library that auto-tags people, places and objects so tags can be searched and browsed, the same scorer and the same tag vocabulary can be deployed in three places. An interviewer asking this is checking whether you can name them and price them — not whether you can describe the model. ## The three placements - **Batch precompute.** A scheduled job walks the stored library — hundreds of millions of photos — scores every row it is asked to and writes a tag row per photo into a store the read path only looks up. The read path never calls the scorer. - **Event-driven scoring.** Each upload publishes an event; a consumer scores that one photo and writes its tag row. The scorer runs once per item, at arrival, off the user's read path. - **Request-time (live) scoring.** No tag row exists until somebody opens the photo. The read that needs the tag calls the scorer synchronously and the viewer waits for it. Batch and event-driven scoring are both *precompute*: the answer is already written when the read arrives. They differ only in the trigger — a clock versus an arrival. Live scoring is the one placement where the scorer runs inside a user-facing deadline. ## What each placement gives and costs | | batch precompute | event-driven | live on read | |---|---|---|---| | a tag first exists | after the next pass reaches the photo | seconds after the upload lands | only after someone opens the photo | | age of a stored tag | up to one cadence | bounded by the consumer's drain time | zero, it is produced now | | scorer calls | one per row per pass | one per uploaded photo | one per distinct photo opened | | coverage of the library | complete after a full pass | complete for arrivals, not for the backlog | only what has been opened | | photos nobody opens | paid for | paid for | not paid for | | scorer inside the read deadline | no | no | yes | | a new scorer version reaches old photos | on the next full pass | only via a separate backfill | on the next open | ## Freshness here means the scorer, not the input "Freshness" is two different clocks and design rounds routinely conflate them: how new the *inputs* are, and how recently the *output* was produced. For viewer-independent tags — objects, places, scene type — the input is the photo's pixels, and those never change after upload. A stored tag for those therefore does not rot with time; it only becomes wrong when the **scorer version** that produced it is superseded. That is why a precomputed tag store is perfectly respectable here and would not be in, say, a risk score over an account's last hour of activity. The exception is a tag whose input genuinely does change: a person tag depends on how the library's owner has clustered and named faces, and that changes long after the batch pass ran. Tags of that kind are the reason pure precompute is often not the whole answer. ## The read pattern usually decides before the cost curve does There are two shapes of read, and only one of them can be served live: 1. **A read that names its keys.** Opening an album hands the service sixty photo ids. A live scorer can produce exactly those sixty predictions, so any placement can serve this read. 2. **A read that ranges over keys the request never names.** A tag search for "beach" must consider the whole stored library. You cannot call the scorer four hundred million times to answer one query, so this read can only be served from rows that already exist. If the product has a search box over tags, the placement argument is already over for the bulk of the corpus: something has to precompute. The cost comparison then decides the *trigger* — clock or arrival — and whether an on-demand path fills the gaps. ## How to answer it in a round 1. Name the three placements and state that two of them are precompute with different triggers. 2. Say which reads the product has, and whether any of them ranges over unnamed keys. 3. Give the cost of each placement in scorer calls per day, using the upload rate and the fraction of photos ever opened. 4. Commit to one, and say what the choice costs you — untagged recent photos, paid-for work nobody reads, or a scorer inside a user-facing deadline.

  • The scorer is too slow to run inside the album-open deadline — what does that fact decide?
    It decides the placement before it decides the scorer. Moving the work off the read path — tagging on the upload event or in a scheduled pass — removes the deadline entirely, so a heavy scorer becomes affordable and its cost is paid asynchronously. Shrinking the work per call so it fits the deadline instead is a different decision, about the millisecond split inside the scoring stage, and it only arises once you have committed to producing the prediction on the request.
  • Does the placement change what the scorer sees as input?
    It can. A batch pass and an event consumer both see only what is stored about the photo. A request-time scorer additionally sees the request itself — who is viewing, on what surface, with what recent activity. If none of that changes the tag, live scoring buys nothing but freshness against the scorer version; if it does, precompute cannot express it at all.
  • Why is a nightly pass over the whole library usually the worst of the three?
    Because it pays the largest number of scorer calls and still gives the weakest freshness guarantee. Rescoring rows whose input has not changed is pure waste, and a photo uploaded just after the pass started waits almost a full cadence for its tags. A scheduled pass earns its place as a backfill for a model change, not as the standing placement.

A kitchen can plate from yesterday's prep list, prep each delivery as it arrives at the door, or cook only what a table actually orders. The food is the same; what differs is how much is thrown away and how long the table waits.

saying these in an interview costs you the question

  • Describes precompute as live scoring with a cache bolted in front
  • Treats freshness as the only axis and never mentions coverage
  • Thinks a nightly pass covers photos uploaded after it started
  • Assumes a tag search can call the scorer once per query
  • Says event-driven tagging removes the need to backfill the existing corpus
  • Claims a stored tag goes stale because the photo itself changed