skip to content

In a candidate send-time model's traffic ramp, what sets how long one step must bake before the next increase?

level: middleimportance: should knowfreq 46%

answer

  1. bake for the signal, not the step
  2. the action has to happen first
  3. one full daily send cycle minimum
  4. users mute later than they receive
  5. queued sends blend the step boundary

basics

~20 s

The latency of the guardrail signals the step is supposed to be halted on. A step must span at least one full cycle of the action being changed plus the users' reaction lag, or the halt it promises can never fire.

solid answer

~40 s

Bake time is a property of the **signal**, not of the step size. A send-time model only acts when sends actually go out, so a step has to cover at least one complete daily send cycle at the new percentage before any candidate behaviour exists to look at. Then the guardrails have their own lag: a user who dislikes a 6am notification mutes the channel later that day or the next, so a window shorter than that lag reads clean by construction. Add settling time — sends queued under the previous percentage are still draining, so the first hours of a step are a blend of both versions. Against that, every extra hour exposes more users and keeps two versions running, which is what caps the window from above.

go deeper

for a junior

Recall that a ramp step is held for a while before the next increase, and that the hold has to be long enough for the change's effects to actually show up.

for a middle

Explain the three delays that set the floor — the action happening, the user reacting, and the boundary settling — and why each pushes the window longer.

for a senior

Show that you derive each step's window from a named signal latency and can say which guardrail is genuinely gating the step and which is only being watched.

for a principal

Own the trade between exposure and evidence across a portfolio of ramps, and set the rule for when a slow guardrail lengthens a step versus being read after the fact.

## What bake time is for The **bake time** of a ramp step is the interval a step is held at its percentage before the ramp is allowed to advance. Its job is narrow and checkable: to make the step's guardrail halt *possible*. A step that ends before its guardrails could have moved has a halt on paper and none in practice, and the ramp then advances on the absence of evidence rather than on evidence of absence. So the question to ask of any proposed schedule is not "is two hours long enough to be safe?" but **"can the signal I would halt on have moved within two hours?"** ## The three delays that set the floor 1. **Exposure delay — the action has to happen.** A send-time model does nothing until the notifications it re-timed are actually sent. If the candidate moves a cohort's notification from 9am to 7am, the effect of that decision does not exist until 7am arrives for those users. A step therefore has to span at least one full daily send cycle, across time zones if the population spans them, before there is any candidate behaviour at all. 2. **Reaction delay — the user responds later than the action.** Opt-out and mute are deliberate acts. A user annoyed by an early notification often mutes at the next convenient moment, not at the moment of delivery. The guardrail's response is therefore smeared over hours after the send, and a window that closes at the send has measured nothing. 3. **Settling delay — the system blends at the boundary.** When the percentage changes, work planned under the old percentage is still in the queue. For the first part of a step the delivered mix is neither the old share nor the new one. Caches, a daily feature refresh and any warmup on the candidate's path add to the same effect. ## What caps it from above | Pressure | Direction | |---|---| | Guardrail signal latency | Pushes the step longer | | Reaction lag of a deliberate user action | Pushes the step longer | | Queue drain and cache warmup at the boundary | Pushes the step longer | | Users exposed to an unproven version | Pushes the step shorter | | Two model versions running and being watched | Pushes the step shorter | | Calendar pressure on the rollout | Pushes the step shorter | Bake time is the point where those two directions balance, and it is legitimately different per step: an early step exposes few users and can be held longer at low cost, while a late step has most of the population on it and the pressure to finish is real. ## Consequences worth stating in a design round - **Guardrails must be observable inside the window you chose.** A signal that only resolves days later cannot gate an hours-long step. Either the step lengthens to match the signal, or the step is gated on a faster-moving signal and the slow one is read after the ramp instead. - **Attribution has to survive the window.** Every send carries the version that produced it, so the guardrail can be computed per side inside the step rather than compared against last week. - **A halt is a hold, not a revert.** Reaching the end of a bake window with a moved guardrail stops the advance; deciding what to do about the exposed cohort is a separate decision. - **Bake time is not the same question as how confidently the effect will eventually be scored.** Here we are asking only whether the signal could have appeared at all inside the window. ## A worked shape For a candidate that changes when a daily notification is sent, a defensible schedule looks like a small first step held across a couple of full send cycles, then a larger step, then larger still — each step at least one complete cycle plus the mute-reaction lag, and each boundary treated as a blend until the previously queued sends have drained. What makes the schedule defensible is not the specific percentages but that each number in it was derived from a delay someone can name.

  • Why can a bake window be shorter for the last step than for the first one?
    The first step exposes few users, so holding it is cheap and its main purpose is to catch a gross failure with a small blast radius. By the last step most of the population is already on the candidate, and the marginal information from holding longer is small while the cost of running two versions and delaying the finish is not.
  • What do you do when the guardrail you care about only resolves days after the send?
    Do not pretend an hours-long step gates it. Either lengthen the step to cover the delay, or gate the step on a faster-moving signal and read the slow one across the whole ramp afterwards, accepting that the ramp advanced without it.

saying these in an interview costs you the question

  • Sets bake time from the step size rather than the signal's latency
  • Ends a step before any affected send has gone out
  • Reads the first hour of a step as pure candidate behaviour
  • Claims a halt exists when the guardrail cannot move in time
  • Uses one fixed bake window for every step regardless of exposure