In a photo library with bursty uploads, what freshness bound can event-driven tagging promise that a nightly tagging job cannot?
answer
- clock trigger versus arrival trigger
- cadence is the worst case, not half of it
- backlog divided by drain rate
- the burst sets the published bound
- at-least-once means idempotent writes
basics
~20 sEvent-driven tagging bounds a tag's age by the consumer's drain time — seconds when the stream is flowing, minutes behind a burst. A scheduled pass can only promise the cadence, so a photo can wait almost a full cycle.
solid answer
~40 sBoth are precompute; the difference is what triggers the scorer and therefore what bound you can put in writing. A nightly pass promises only its cadence: a photo uploaded just after the job starts waits nearly 24 hours. An arrival-triggered consumer promises drain time — the queued backlog divided by the consumer's throughput. When uploads flow evenly that is seconds. Under a bulk import of 50,000 photos into a consumer draining 500 items per second it is about 100 seconds, and the promise you publish has to be the burst figure, not the quiet one. What the queue buys is a *bounded* delay instead of an unbounded one; what it costs is idempotency work, because a redelivered upload event must not write a second tag row.
go deeper
Know that both are precompute and that the difference is what starts the scorer: a schedule or the arrival of a photo. The scheduled one leaves recent photos untagged for a while.
Compute a drain time from a backlog and a consumer rate, and state a nightly pass's worst case as a full cadence rather than half of it. Say why a queue converts a burst into a bounded lag.
Publish a bound you can hold under the worst arrival burst the product allows, and design the consumer for redelivery, poison items and events that outrun their bytes. Keep the scheduled pass for backlog and version rescores.
Decide what freshness is worth as a product promise before buying drain rate to hit it. A tighter bound for ordinary uploads and a looser one for bulk imports is often a better bargain than a single aggressive number.
## Two triggers inside one placement family A nightly pass and an upload-event consumer are both precompute: the tag row exists before the read arrives, and the read is a lookup either way. They differ only in what starts the scorer — a clock or an arrival — and that single difference decides the freshness bound you can promise, the shape of the load you must size for, and the operational bill you sign. ## What a scheduled pass can promise A pass that runs once a day can promise exactly one thing: *no photo is older than one cadence without tags, measured from the pass that follows it*. The worst case is a photo uploaded seconds after the job began scanning — it misses this run entirely and waits for the next. In practice that means a promise of "up to about 24 hours", not "about 12". What the schedule buys is predictability in the other direction. The work is known in advance, it is bounded by the corpus, and it can run on whatever capacity is cheapest at that hour. Nothing about a bulk import at midday changes the shape of the job. ## What an arrival-triggered consumer can promise An event consumer's freshness bound is the **drain time** of its backlog: ``` lag_seconds = queued_items / consumer_throughput_per_second ``` With an evenly flowing stream, the queue is near-empty and the bound is dominated by the scoring call itself — seconds. The interesting case is the burst. Consumer photo uploads are not a smooth Poisson arrival process: they are diurnal, and a single owner syncing a camera roll can drop tens of thousands of photos into the stream in under a minute. A 50,000-photo import hitting a consumer that scores 500 photos per second drains in 100 seconds. If you have published "tags within 10 seconds", that import has already broken the promise. ## The queue is what makes the bound bounded Without a buffer, arrival-triggered scoring means the scoring fleet must be sized for the peak arrival rate or drop work. The queue converts the burst into a backlog, which changes the guarantee's shape rather than its existence: - **Without a queue:** the promise is "instant, until the fleet saturates, then errors". - **With a queue:** the promise is "within the drain time", which degrades smoothly and is a number you can publish, alarm on and pay to shrink. The consumer throughput you buy is then a straightforward trade: doubling it halves the worst-case lag under the same burst, and the arithmetic is visible to everyone. ## The operational bill event scoring adds A scheduled pass over a stored library is idempotent by construction — rerun it and it writes the same rows. An event consumer is not, and that difference is where the real work is: 1. **Redelivery.** At-least-once delivery means the same upload event can arrive twice. The write has to be keyed by the photo so a second delivery overwrites rather than duplicates. 2. **Poison items.** One photo the scorer cannot handle must not stall the partition behind it; it needs a bounded retry and then a side channel. 3. **Ordering against the upload.** The event can outrun the stored bytes it refers to, so the consumer must handle "not readable yet" as a retry rather than a failure. 4. **The existing backlog.** Events only cover arrivals. Everything stored before the consumer existed needs one scheduled pass, which is why the two triggers usually coexist rather than compete. ## Choosing between them | | nightly pass | upload-event consumer | |---|---|---| | freshness bound | the cadence, worst case a full cycle | the queue's drain time under peak | | load shape | known, bounded, schedulable off-peak | follows the arrival process, bursty | | work per photo | one call per pass | one call per photo, ever | | covers the stored backlog | yes | no, needs one pass | | idempotency | free, the pass is repeatable | must be engineered | The usual production answer is not one or the other. A standing consumer on the upload stream gives new photos tags in seconds, and a scheduled pass exists for exactly two jobs: covering everything uploaded before the consumer, and rescoring the corpus when a new scorer version ships. If you find yourself running a full nightly pass as the *standing* placement, the question to ask is what it is producing that the previous night's run did not.
- How do you keep a redelivered upload event from writing a duplicate tag row?Key the write by the photo id and the scorer version and make it a replace rather than an append, so a second delivery converges on the same row. The consumer stays stateless and correctness does not depend on exactly-once delivery, which the stream is unlikely to give you anyway.
- A bulk import blows the published freshness bound. What do you change?Either buy drain rate so the burst clears inside the bound, or publish a bound that is honest about imports — for example a tighter promise for ordinary uploads and a looser one for a sync of more than a few thousand photos. Quietly missing the number is the one option that is not available.
- Is there any reason to keep a nightly pass once event tagging works?Two. It covers everything uploaded before the consumer existed, and it is the vehicle for rescoring the corpus when a new scorer version ships. Both are finite jobs rather than a standing placement, so the pass runs on demand rather than every night.
saying these in an interview costs you the question
- Claims event-driven tagging is instant regardless of arrival rate
- Sizes the consumer for the average upload rate and ignores imports
- Says a queue removes the freshness bound rather than bounding it
- Quotes half the cadence as a nightly pass's worst-case tag age
- Assumes a full pass can simply run hourly at the same cost
- Forgets a redelivered upload event must not create a second tag row