skip to content

Why can a candidate send-time model ramped to 10% of users deliver two notifications to one user in a day?

level: middleimportance: must knowfreq 62%

answer

  1. same user, same side
  2. assignment is a function, not a draw
  3. two planners, one user, one day
  4. stamp the version on the queued send
  5. claim a key of user, date, campaign

basics

~20 s

Because the ramp assignment was not sticky. The same user fell on the candidate side in one planning run and on the incumbent side in another, so both sides queued a send for that day. Deterministic per-user assignment with one claimed send key prevents it.

solid answer

~40 s

A model's prediction here is not a read — it causes an external action, so an unstable assignment produces a duplicate side effect rather than a slightly noisy measurement. The fix has two halves. First, the side is a *function* of the user: a deterministic hash of a stable user id plus a fixed ramp identifier into buckets, compared against the current percentage. Raising the percentage only moves users from the incumbent side to the candidate side, never back, so nobody flaps. Second, the decision is resolved once per send window and **stamped onto the queued send** together with an idempotency key of user, send date and campaign, which the sender claims before enqueueing. Assignment is never re-evaluated for sends that are already queued; a percentage change applies to the next window.

code

pseudocode · 10 lines
pseudocode
bucket = hash(user_id + ":" + ramp_id) mod 10000
side   = if bucket < ramp_percent * 100 then "candidate" else "incumbent"

key = user_id + ":" + send_date + ":" + campaign_id
if not send_log.claim(key):
    skip user                      // another planner already claimed this day

model = load_version(side)         // exactly one side scores this user
hour  = model.predict_hour(user_id)
enqueue(user_id, hour, side, model.version, key)

go deeper

for a junior

Remember that the side of a ramp must come out the same every time for the same user, and that it is computed from a stable user identifier.

for a middle

Explain the bucket-and-threshold mechanism, why the threshold only rises, and why the resolved side is stamped onto the queued send instead of recomputed later.

for a senior

Diagnose the duplicate send as two writers both claiming a user, and reach for a claimed idempotency key of user, send date and campaign as the structural fix.

for a principal

Decide the assignment unit across the organisation — person, account or household — because a model whose output is an external action inherits whatever unit the ramp chose.

## Why a duplicate send is the failure, not a noisy number In a notification service, the model's output is acted on: the predicted hour becomes a real send. That makes the ramp's assignment an **operational** invariant, not only a measurement one. If a user can be resolved to the candidate side at one moment and the incumbent side at another, two different code paths each believe they are responsible for that user's notification, and the user receives two of them. The user does not see a slightly different send hour; the user sees spam, and the guardrail you are watching — opt-out or mute rate — moves for a reason that has nothing to do with whether the candidate picks good hours. ## The two properties the assignment needs **1. It must be a function, not a draw.** The side is computed from inputs that do not change during the ramp: - a **stable user identifier**, not a session, device or request id; - a **fixed ramp identifier** mixed in, so this ramp's split is independent of any other split running at the same time; - a bucket space large enough to express the percentages you intend to use. Because the same inputs always produce the same bucket, no state has to be stored, no lookup can go stale, and a planner restarting mid-run resumes with the same answers. Crucially, the comparison is `bucket < threshold`, and the threshold only *rises* as the ramp advances — so the only movement is **incumbent to candidate**. A user never returns to the incumbent side because the percentage changed, which is what would otherwise make a user's send hour oscillate day to day. **2. The decision must be resolved once and carried.** The planning job resolves the side, then writes it onto the queued send along with an **idempotency key** — user id, send date, campaign — that is claimed before anything is enqueued. Anything downstream reads the stamped version rather than recomputing it. This is what makes the assignment survive: - a planner retry after a partial failure; - a ramp percentage change made while a day's sends are already planned; - two planning jobs running concurrently during a deploy overlap. ## How the duplicate actually arises | Mistake | What the user experiences | |---|---| | Side drawn randomly per scoring pass | Two planners both claim the user; two notifications | | Hash includes a run id or timestamp | The bucket changes between runs, same outcome | | Percentage re-applied to already-queued sends | A send planned by the incumbent is re-planned by the candidate | | Two planners with separate queues and no shared claim | Each is internally consistent, and together they double-send | | Assignment keyed on device rather than user | A multi-device user is on both sides at once | The last row is the subtle one: the unit of assignment has to be the entity the **action** is aimed at. A notification is aimed at a person, so the person is the unit. ## The single-writer discipline Even with a perfect hash, two processes can both decide to send if neither is authoritative. The durable fix is that exactly one writer claims the right to send to a user for a given day. A claim on the idempotency key is that writer's ticket: the first claim wins, the second planner sees the key taken and skips the user entirely. This costs one conditional write per user per day and removes the whole class. ## What this does *not* solve Stickiness keeps a user on one side; it does not make the side's behaviour good. A candidate that predicts a 5am hour will send at 5am, once, to exactly the 10% assigned to it — and that is precisely the point. A correct ramp makes the harm *proportional and attributable*: one send per user per day, from a known version, to a known share of the population. Everything else in the rollout — guardrail halts, the bake window, the revert — depends on that attribution being true.

  • When the ramp rises from 10% to 25%, why does no user need a new assignment key?
    The bucket is a fixed function of the user id and the ramp id, and only the threshold it is compared against moves. Raising the percentage therefore moves users from the incumbent side to the candidate side and never back, so nobody flaps between two send hours as the ramp advances.
  • If the candidate were a ranking model rather than a send-time model, would the assignment unit change?
    No — assignment stays per user. One user's whole session is then served by one ranker version, so the list they see is internally consistent and every impression can be attributed to a side. Splitting per request would mix two versions inside a single list and make both the behaviour and its attribution incoherent.
  • What should happen to sends already queued under the candidate when the ramp percentage is lowered?
    They keep the version stamped on them, or they are explicitly drained and re-planned as one deliberate action. What must never happen is a silent re-evaluation of assignment for queued work: that is exactly the path that produces a duplicate send or a dropped one.

saying these in an interview costs you the question

  • Draws the ramp side randomly at each scoring pass
  • Treats the harm as a skewed readout rather than a duplicate send
  • Hashes the user id together with the run timestamp
  • Re-evaluates assignment for sends already queued
  • Assumes two planning jobs writing sends need no shared claim
  • Keys the assignment on a device or session instead of the user