skip to content

You need a feature-toggle system used by 100+ microservices to gradually roll out a new checkout flow. What toggle types would you consider, where should toggle state live, and what should happen to a request if the toggle-evaluation service is unavailable mid-request?

level: seniorimportance: must knowfreq 60%

answer

  1. release vs ops vs permission toggles
  2. local in-memory cache, not per-request network call
  3. decouple deploy from release
  4. fail open to safe default
  5. toggle debt -> delete promptly

basics

~20 s

Use an on/off switch that can target a percentage of users, store its state somewhere fast and locally readable rather than deep inside one remote service, and make sure that if the switch-checking system goes down, requests fall back to a safe default instead of failing or exposing the risky new feature to everyone.

solid answer

~40 s

Feature toggles for a gradual rollout are typically 'release toggles' — short-lived, meant to be deleted once the flow is fully shipped — distinct from longer-lived 'ops toggles' (kill switches) or 'permission toggles' (entitlements). State should live in a fast, low-latency local cache the client SDK reads, refreshed via polling or streaming from a central flag service, rather than a synchronous network call per request, since that adds latency and a new failure dependency to every request path. If the flag-evaluation service is unreachable, the client should fall back to the last-known cached value or a safe hardcoded default — fail open to the old, proven checkout flow — rather than failing the request or defaulting to the new, less-proven path.

go deeper

for a junior

Should know a feature toggle is basically an if/else controlled by a config value that turns a feature on or off.

for a middle

Should know toggles let you deploy without releasing, and can describe percentage rollouts at a basic level.

for a senior

Should design the local-cache/SDK evaluation model to avoid a per-request dependency, and reason about the fail-open default and toggle debt.

for a principal

Should weigh platform-wide governance of toggle lifecycle across 100+ services and teams — ownership, expiry policy, and the org-level risk of a shared toggle-service outage as a single point of failure.

## What a toggle is, and which kind this is Feature toggles are conditionals in code — `if (flags.isEnabled("new-checkout", context)) { ... } else { ... }` — that let you deploy code and release behavior as two separate events. For a gradual checkout rollout you'd primarily use a **release toggle**: short-lived, meant to exist only for the rollout window, evaluated per request with a targeting rule such as a percentage rollout or a consistent-hash bucket so the same user always lands on the same variant. This differs from - an **ops toggle**, a longer-lived kill switch to disable a feature under load; - an **experiment toggle**, A/B bucketing tied to analytics; - and a **permission toggle**, an entitlement gate such as only enterprise-tier customers seeing a feature. Conflating these types in one system tends to create **toggle debt**, because they have different lifecycles — release toggles should be deleted within weeks, permission toggles live indefinitely. ## Where toggle state should live At a scale of 100+ services, a naive design where every request makes a synchronous network call to a central toggle service to ask "is this on?" adds - a new network hop, - a new latency tail, - and a new hard dependency to every request path platform-wide — if that service has a bad day, everything using toggles degrades or fails together. The standard fix is a **pull-and-cache** model: each service embeds an SDK that polls the central flag service, or subscribes to a streaming update, on a short interval and caches the full flag ruleset in memory locally; evaluating "is this flag on for this user" then happens entirely in-process, with zero network calls in the request's hot path. ## Why it exists This exists to **decouple deploy from release**: you can merge and deploy the new checkout code dark, defaulted off, days before turning it on, de-risking the deploy itself, then ramp exposure from 1% to 10% to 100% independently of any deploy, and instantly roll back exposure by flipping the toggle off, with no redeploy, if metrics regress — far faster than a code rollback and redeploy cycle, which is the whole point of separating "ship the code" from "expose the behavior." ## The cost, in both directions The cost runs both ways. - **Every toggle multiplies the code paths that must be tested** — the old flow, the new flow, and the transition state where different requests hit each — and toggles left in code long after rollout completes become dead, confusing branches, so teams need discipline or automated linting to delete release toggles promptly. - **The local-cache model trades perfect real-time consistency for availability and latency.** There's a propagation delay equal to the poll interval where different service instances can briefly disagree about a flag's state, so toggle-dependent logic must tolerate a short window of fleet-wide inconsistency rather than treat toggle checks as an atomic global truth. ## What happens when the toggle service is down If the toggle-evaluation service is unavailable and the client SDK isn't built defensively, the naive failure is 1. the request throwing, if the call is synchronous and unguarded, 2. or defaulting to whatever value is coded as the SDK's "unreachable" fallback — and if that default is misconfigured to the new, less-proven checkout path, an outage of the toggle service can accidentally roll out an unvetted feature to 100% of traffic at the worst possible time. The safe design is **cache-first evaluation** with no per-request network call at all, and if the cache itself is empty or stale beyond a threshold, **failing open** to the known-safe old behavior, with the client library shipping a mandatory default value per flag as a deliberate safety net. ## Where it shows up This is roughly how LaunchDarkly, Unleash, and Split.io are architected: an SDK in each service process streams or polls flag rule updates, evaluates flags fully client-side, and requires a mandatory default-value parameter on every evaluation call precisely so a network partition to the flag service degrades to a known value instead of an exception. Companies like Etsy and Flickr popularized this "flags for continuous deployment" pattern specifically to decouple risky rollouts, like a checkout rewrite, from the deploy pipeline, ramping exposure gradually while watching conversion rate and error rate at each step before widening further.

  • How would you implement a 'percentage rollout' so the same user consistently sees the same variant across multiple requests, instead of flickering between old and new?
    Use consistent hashing: hash a stable identifier like the user ID combined with the flag's name into a fixed range, say 0-99, and compare it against the rollout percentage threshold — the same user ID always hashes to the same bucket for that flag, so they deterministically land in the same variant on every request without storing per-user assignment state anywhere. This also lets you widen the percentage later, from 10% to 25% say, while guaranteeing everyone already included stays included, since the hash-to-bucket mapping never changes.
  • What's toggle debt, and how would you prevent it at a platform-wide scale with 100+ services?
    Toggle debt is the accumulation of stale, no-longer-needed conditional branches left in code after a rollout completes, making the codebase harder to reason about and potentially hiding latent bugs in the untested, abandoned branch. Prevention typically means requiring every release toggle to have an owner and an expiry date at creation, with automated reports flagging toggles older than some threshold for cleanup, and making toggle removal part of the definition-of-done for the rollout.
  • Why might you use a consistent-hash-based rollout instead of simply storing 'flag enabled: true/false' per user in a database?
    Storing per-user state requires a database read on every evaluation and a write the first time each user is bucketed, adding load and a stateful dependency across potentially millions of users and 100+ services; a hash-based bucket is computed in-memory with zero storage and zero I/O, and is naturally consistent across every service without any synchronization. The trade-off is giving up the ability to arbitrarily and permanently override one specific user's bucket without layering a separate explicit targeting rule, like an allowlist, on top.

It's like a dimmer switch wired to a local relay in the room rather than a call center you phone every time you flip a light — the relay periodically syncs its setting from headquarters, but flipping the light itself never depends on the phone line being up.

saying these in an interview costs you the question

  • Proposes a synchronous network call to a central toggle service on every request
  • Treats all toggle types (release, ops, permission) as the same with no lifecycle distinction
  • Doesn't have an answer for what happens when the toggle service is down
  • No mention of consistent bucketing so users don't flicker between variants
  • Leaves toggles in code indefinitely with no cleanup plan

context