While a candidate send-time model serves 10% of users, why does the previous model version stay loaded and serving too?
answer
- a split, not a swap
- both versions serve at once
- the other 90% still needs an owner
- revert becomes a routing change
- incumbent stays warm as the comparison side
basics
~20 sA percentage ramp splits live traffic between two model versions instead of replacing one with the other. The previous version keeps serving the other 90%, and keeping it loaded is what makes reverting a routing change rather than a redeploy.
solid answer
~50 sA ramp is a split, not a swap. In a notification service that predicts each user's best send hour, a 10% step means one user in ten gets the hour the candidate model version predicts and nine in ten get the hour the previous model version predicts, in the same planning run, over the same user base. So both model artifacts are loaded and both feature paths are live for the whole ramp. That buys three things: somebody is still serving the majority, the previous version is the side the guardrails compare the candidate against within the same window, and reverting exposure is a change to the ramp percentage rather than a rebuild and redeploy. The price is a temporary one — two artifacts in memory, a wider config surface, and a ramp that has to actually finish.
go deeper
Recall that a percentage ramp splits users between two model versions that are both running, and that the previous version is still serving the majority.
Explain why the incumbent staying warm is what turns a revert into a routing change, and name what the second resident artifact actually costs the fleet.
Show that you design the ramp so the state you want to return to is already running, and that every served notification records which model version produced it.
Frame the trade: each concurrently ramped version adds memory, config surface and an unexplained-behaviour class, so the number of simultaneous ramps is a budget the organisation sets, not a per-team choice.
## What a percentage ramp actually does A **candidate model version** is a newly trained artifact that has already cleared whatever bar the team set before any user sees its output. A **traffic ramp** is the stage after that bar: the candidate is given a small, growing share of real users, and its predictions are acted on for those users. The worked system here is a notification service whose job is to pick, per user, the hour at which today's notification goes out. At a 10% ramp step, the planning job resolves each user to one side and then asks that side's model version for an hour. The critical point for a first design round is that **nothing is removed when a ramp starts**. The **incumbent model version** — the one already in production — is still deployed, still loaded, still scoring, and still responsible for the majority of users. The system is running two model versions at once on purpose, and that condition lasts until the ramp either completes or is reverted. ## Why the incumbent has to stay live 1. **It owns the traffic the candidate does not.** At 10%, 90% of users still need a send hour. There is no third thing serving them; they are served by the previous model version exactly as they were the day before. 2. **It is the comparison side.** A guardrail halt is only meaningful when the candidate's behaviour is read against the incumbent's over the *same* window, on the *same* kind of population. If the incumbent were switched off, the only comparison left would be against last week, which mixes the model change with everything else that changed. 3. **It makes reverting cheap.** When the previous version is already warm, undoing exposure is a change to one number — the ramp percentage — or a switch that routes everyone back. If instead the previous artifact had been unloaded, undoing exposure would mean a build-and-deploy cycle, and recovery time would be measured in deploy minutes rather than in config seconds. 4. **It is the sane fallback.** If the candidate's scoring path fails for a user — a missing feature, a timeout in the planning job — the previous version is right there to produce an hour, so the user still gets a notification at a defensible time. ## What running two versions costs | Resource | What the ramp does to it | |---|---| | Model artifacts in memory | Doubles: both versions are resident on the serving or planning fleet | | Per-user scoring work | Roughly unchanged: each user is scored by one side, not both | | Feature reads | One path per user, plus any additional feature the candidate needs | | Config surface | Grows: ramp percentage, ramp identifier, halt state, pinned versions per side | | Operational attention | Grows: two behaviours to explain when someone reports a badly timed notification | The cost is real but bounded, and it is the reason a ramp is a *temporary* state with a schedule and an end, not a permanent architecture. A team that leaves three model versions ramped indefinitely has three behaviours in production and no clear owner for any user complaint. ## The shape of one step - **Assign.** Each user is resolved to the candidate side or the incumbent side, and that resolution has to be stable for the user across the whole ramp. - **Serve.** The assigned side produces the send hour; the resulting notification carries the version that produced it, so every downstream measurement can be attributed to a side. - **Bake.** The step is held at its percentage long enough for the guardrails to be able to move at all. - **Read.** Guardrails are compared per side inside the step's window. - **Decide.** Raise to the next percentage, hold where you are, or revert. ## What the ramp is and is not A ramp is the first stage at which a user actually *receives* the candidate's decision, so it is also the first stage that can do harm. That is precisely why the incumbent is kept warm underneath it: the entire value of a staged rollout is that the state you want to return to is already running. A rollout design where "go back" requires rebuilding something is not a staged rollout; it is a deploy with extra steps.
- An offline head-to-head picked the candidate as the winner. Does it go straight to 100% of users?No. Winning an offline comparison only qualifies a candidate to *start* the ramp. It then earns traffic in percentage steps, each held for a bake window with guardrails that can halt it, while the incumbent keeps serving everyone else. The offline result decides whether the ramp begins, not how much traffic the candidate gets.
- What does the fleet pay for keeping both model versions live?Memory for two resident artifacts, warm caches for both, and any extra feature the candidate needs materialised alongside the ones the incumbent reads. Per-user scoring work barely changes, because each user is scored by one side. The bigger cost is config and attention: two behaviours in production that both have to be explainable.
saying these in an interview costs you the question
- Thinks a ramp replaces the previous model version everywhere
- Assumes reverting exposure means redeploying the previous artifact
- Describes the unramped majority as unserved or served by a static default
- Treats running two model versions as a permanent architecture
- Cannot say which side a given user's notification came from