When a feature flag system does a percentage rollout, say enabling a flag for 10% of users, how does it decide which specific users fall in that 10%, and why must that decision be consistent across repeated evaluations for the same user rather than random each time?
answer
- hash(flagKey+userId) -> bucket
- sticky = deterministic per user
- monotonic: raising % only adds users
- flag-level % rollout vs infra canary deploy
- low % = noisy sample size
basics
~20 sThe system turns each user's ID into a number (via a hash) and checks if that number falls under the 10% cutoff. Because the same ID always hashes to the same number, the same user always lands on the same side of the line, so they don't flicker between old and new experience.
solid answer
~40 sA percentage rollout typically hashes a stable identifier (user ID, device ID, or session key) combined with the flag's own key into a number in a fixed range (e.g., 0-9999), then compares it against a threshold (e.g., under 1000 means 'in' for 10%). Because the hash is deterministic, the same user always gets the same bucket for that flag -- this is sticky bucketing -- so raising the rollout from 10% to 20% only needs to add users, not reshuffle everyone, and a given user's experience is stable across requests and sessions. Without this, evaluating fresh randomness per request would cause the same user to flicker between variants, breaking any multi-step flow and making bugs impossible to reproduce consistently.
go deeper
Should get that the same user needs to consistently land in or out of a rollout, even if they can't explain hashing mechanics.
Should describe hashing a stable user identifier into a bucket and explain why that gives sticky, per-user-consistent behavior.
Should explain the monotonic-growth property, the flag-percentage vs infrastructure-canary distinction, and at least one concrete failure mode like cross-SDK hash mismatch or unstable identifiers.
Should discuss statistical sample-size implications at low percentages, monitoring segmentation by bucket during a canary, and how this pattern composes with targeting rules and infra-level canary deploys in a real release process.
## How a user lands inside the slice Mechanically, a percentage rollout answers one question per request: does this particular user fall inside the enabled slice? The standard technique is **deterministic hashing** rather than a coin flip. Step by step: 1. The flag system takes a **stable identifier** for the user -- a user ID, device ID, or session key. 2. It combines that with the flag's own key, so the same user gets an independent-looking assignment for every different flag. 3. It runs the pair through a fast, well-distributed hash function. 4. That hash is mapped onto a fixed numeric range, commonly 0-99 or a finer 0-99999 for more precision at low percentages, producing what's usually called the user's **bucket** for that flag. Enabling the flag for 10% of users then just means: if bucket is under 10 (or under 10000 out of 100000), evaluate to true. Because hashing a given flag-key-and-user-id pair always produces the same bucket number, the same user gets the same answer every single time they're evaluated, on any server, in any request -- that's what **sticky bucketing** means. ## Why stickiness is the whole point Stickiness is not a nice-to-have; it's the entire point. Two properties fall out of it that would otherwise be impossible. 1. **First, consistency within a user's experience.** If the flag gates a multi-step flow such as a new checkout form, the user needs to see the same variant on step one and step five, or the flow breaks or looks buggy; a re-randomized-per-request assignment would flicker the user between old and new UI on every page load. 2. **Second, monotonic growth -- just as important operationally.** Because the bucket number for a user never changes, raising the rollout threshold from 10% to 20% is guaranteed to be a superset -- every user already in at 10% stays in, and only new users are added -- rather than reshuffling the entire population. That monotonic property is exactly what a staged canary rollout depends on: engineers watch error rates and latency at 1%, then 5%, then 25%, then 100%, confident that the users already exposed keep being exposed and the blast radius only grows in one direction. ## Why the pattern exists: risk control The reason this pattern exists at all is risk control during release. Flipping a flag to 100% for every user simultaneously means any bug, performance regression, or unexpected load pattern hits your whole user base at once, with no early warning. A percentage or canary rollout turns a release into a controlled experiment against production traffic: start small, watch dashboards and alerts, and only widen exposure once the smaller cohort looks healthy. It's worth distinguishing this flag-level canary from an infrastructure-level canary deployment -- the latter is about validating the new build itself. | Flag-based percentage rollout | Infrastructure-level canary deployment | |---|---| | A flag-based rollout instead varies behavior within a single, already-deployed binary | Routes a small percentage of traffic to a new binary or version running on a subset of instances | | Based on which users are bucketed in | Independent of any user identity | The two are complementary and often layered: you might canary a new binary to 5% of instances first, and separately, within that binary, still gate the risky feature behind its own flag rollout percentage. ## The trade-offs The trade-offs run in both directions. - Sticky bucketing gives consistency and monotonic rollout, but it means you cannot cleanly do a targeted, per-user rollback: if one specific unlucky user in the 10% cohort is hitting a bug, you can't surgically exclude just them without a separate explicit exclusion rule layered on top; the only blunt instrument is rolling the whole percentage back, which removes everyone. - At low percentages, hashing also introduces statistical noise. 1% of a million users is a solid sample, but 1% of ten thousand users is a much noisier signal, so error-rate comparisons between in and out cohorts need enough absolute users to be meaningful. ## Failure modes The characteristic failure modes are worth naming concretely. 1. If the identifier used for hashing isn't actually stable -- for example bucketing by a request-scoped correlation ID for a user who is sometimes anonymous and sometimes logged in -- the same human gets reshuffled into a different bucket on every visit, silently defeating stickiness and reproducing the flicker problem described above. 2. Another common failure is when client-side and server-side SDKs for the same flag platform implement the hash differently or seed it differently, so a mobile app and a backend service disagree about which bucket a given user is in, producing an inconsistent cross-surface experience. 3. A third is forgetting to segment monitoring by bucket during a canary: if error-rate dashboards aggregate all traffic together, a real regression hiding inside the 5% in cohort gets diluted into noise by the 95% unaffected cohort and never trips an alert. ## Where it shows up in the wild A concrete real-world instance: LaunchDarkly's percentage rollout implementation is exactly this deterministic-hash-into-bucket model, and it explicitly documents that rollouts are sticky per user and that raising a rollout percentage only adds users, never reassigns existing ones -- precisely so staged canary releases behave the way on-call engineers expect when they widen exposure step by step while watching dashboards.
- Why bucket on hash(flagKey + userId) instead of just hash(userId) alone?Hashing only the user ID would make a given user land in the same relative position for every single flag in the system, correlating exposure across unrelated rollouts. Mixing in the flag's own key makes each flag's bucketing independent, so a user being in the 10% for one feature has no bearing on whether they're in the 10% for another.
- A team raises a rollout from 10% to 5% to roll back a partially-bad release, then plans to go back up to 10% once fixed. What goes wrong with sticky bucketing here?Lowering the percentage isn't monotonic the way raising it is -- some users who were previously in at 10% will now fall outside the 5% cutoff and get pulled back to the old behavior, while others stay in, so the rollback is a real behavior change for a subset of users, not a no-op. Going back up to 10% later re-adds exactly the same users as before since the hash is deterministic, but the down-then-up cycle still causes visible flicker for users near the boundary.
- If a mobile app's SDK and the backend API's SDK for the same flag platform hash bucket assignment slightly differently, what production symptom would that cause?A given user could be bucketed in the rollout on one surface, say the app shows new UI, but out on another, say the backend still serves old-behavior responses, producing a broken or inconsistent experience. This is why flag SDKs from the same vendor guarantee identical hashing algorithms across all their language SDKs.
Like a raffle where your ticket number is computed from your name, not drawn fresh each time -- you always land in the same numbered slot, so raising the winning-number cutoff from the first 10 numbers to the first 20 numbers only adds new winners, it never un-wins someone who already won.
saying these in an interview costs you the question
- Thinks percentage rollout means a fresh random coin flip on every request
- Doesn't know raising the rollout percentage should be monotonic (superset of prior users)
- Can't explain why the same user needs a consistent flag value across requests
- Conflates flag-based percentage rollout with infrastructure canary deployment as the same mechanism
- Assumes 1% is always a statistically meaningful sample regardless of total user count