Feature flags are often grouped into categories such as release flags, ops or kill-switch flags, experiment flags, and permission flags. What distinguishes these categories, and why does each need a different lifecycle or cleanup policy?
answer
- release=temporary
- ops=permanent safety valve
- experiment=sticky bucketed A/B
- permission=business logic not debt
- different owners, different expiry
basics
~20 sNot all flags are used the same way: some hide a feature until it's ready (release), some let you turn something off in an emergency (ops), some split traffic to compare two versions (experiment), and some control who is allowed to use a feature long-term (permission). Because they serve different purposes, some should be deleted quickly and others are meant to live forever.
solid answer
~50 sRelease flags are short-lived and binary, meant to be deleted within days or weeks of 100% rollout since their only job is decoupling deploy from release. Ops/kill-switch flags are long-lived safety valves around risky dependencies (a third-party API, a cache) that stay in the code indefinitely so on-call can disable a failing subsystem instantly. Experiment flags drive A/B tests, are often multivariate, tied to an analytics pipeline, and are removed as soon as the experiment concludes and a winner is picked. Permission/entitlement flags gate features by plan tier or user segment (e.g., enterprise-only SSO) and are effectively permanent business logic, not technical debt. Treating them all the same -- e.g., trying to clean up a permission flag, or letting a release flag live for a year -- is the core mistake; each category needs its own owner and expiry policy.
go deeper
Should recognize that not all flags are the same and give a rough example of two different purposes, e.g. 'turn a feature on' vs 'turn something off in an emergency.'
Should name and correctly distinguish at least three of the four categories and know release flags should be deleted after rollout while ops flags shouldn't.
Should explain the different lifecycle and ownership implications for each category and the risk of applying one category's cleanup policy to another.
Should discuss this as an organizational/tooling problem -- how flag inventory systems and process need to encode category to avoid both flag rot and accidental deletion of permanent flags at scale.
## Why intent, not mechanism, is the dividing line Feature flags are a single mechanism -- a runtime switch controlling a code path -- but the industry has converged on classifying them by intent, because intent determines lifespan, who owns the decision to flip them, and how dangerous it is to forget about them. The four commonly named categories are release flags, ops (kill-switch) flags, experiment flags, and permission (entitlement) flags. | Category | Lifespan | What it is for | |---|---|---| | Release | Temporary by design | Implements the deploy/release decoupling | | Ops (kill switch) | Meant to live indefinitely | Functionality with an identified failure mode | | Experiment | Bounded by statistical significance | Two or more variants of behavior against a metric | | Permission (entitlement) | Semantically permanent | Real business logic | ## The four categories in detail **Release flags** exist purely to implement the deploy/release decoupling: wrap an in-progress feature so it can merge to trunk and ship dark, then flip on for real users when it's ready, often via a staged percentage rollout. Their defining property is that they are temporary by design -- once the feature reaches 100% rollout and has proven stable for some bake-in period, the flag has done its job and both the flag check and the old code path it replaced should be deleted. A release flag still in the codebase six months after full rollout is pure liability: it adds a branch nobody tests anymore and signals to new engineers "this might still be turned off," which is false. **Ops flags (kill switches)** are the opposite in lifespan: they are meant to live indefinitely. They wrap functionality with an identified failure mode -- a call to a flaky third-party payment processor, a new caching layer, a recommendation engine -- so on-call can disable just that piece in seconds during an incident without touching anything else. Removing an ops flag is itself a risk decision, not a cleanup task, because its value is precisely in being there, ready the one time it's needed. The design question for ops flags is less "when do we delete this" and more "what happens when it's flipped off -- is there a safe degraded fallback rather than an error?" **Experiment flags** (A/B or multivariate test flags) exist to compare two or more variants of behavior against a metric such as conversion rate or click-through, usually by randomly and consistently (sticky, hashed by user ID) assigning users to a bucket and piping the assignment into an analytics platform. Unlike release flags, they're often not binary; they may have several named variants such as control, treatment-A, treatment-B. Their lifecycle is bounded by statistical significance: once the experiment reaches its pre-registered sample size or duration and a winner is declared, the flag should collapse -- the winning variant becomes the permanent code path, losing variants get deleted, and the flag itself is removed. Leaving an experiment flag live past its conclusion both bloats the code and can silently keep splitting traffic in a way that skews future experiments or dashboards. **Permission (entitlement) flags** are structurally identical to a release flag -- an if/else -- but semantically permanent: they encode real business logic like "SSO is only available on the Enterprise plan" or "this admin tool is only visible to internal staff." These are not technical debt to be cleaned up; they're first-class product configuration that happens to be implemented with the same flag infrastructure, and treating them like release flags would be a functional bug, removing paid-tier gating from the product. ## Why tooling and process must key off the category The reason a single taxonomy matters in practice is that flag-management tooling and process both need to key off category. A flag inventory dashboard can flag a release-type toggle that's been at 100% for sixty days as stale and worth removing, while never flagging an ops or permission flag the same way, because their staleness is normal, even desirable. Getting this wrong causes two opposite failure modes in production: - **Release-flag rot**, where nobody ever prunes short-lived flags and the codebase accumulates hundreds of dead branches nobody dares delete because it's unclear which category they belong to. - **Accidental deletion of a permanent flag** during a flag-cleanup sprint that mistook it for a stale release flag, silently breaking a kill switch or removing paywall gating. ## Where the split shows up at scale A concrete real-world illustration: large-scale platforms like Netflix run their internal experimentation tooling and kill-switch/circuit-breaker tooling as explicitly separate systems from ordinary release-flagging tooling for exactly this reason -- the operational and statistical guarantees each category needs (instant global propagation for kill switches, sticky consistent bucketing with analytics integration for experiments) are different enough that conflating them, even though the underlying runtime-conditional mechanism is identical, would create both engineering and organizational confusion about who owns what and when it's safe to remove.
- Why is it dangerous to run a 'delete all flags older than 90 days' cleanup script across an entire flag inventory without regard to category?It would delete ops/kill-switch and permission flags that are supposed to be old and permanent, silently removing an incident safety valve or paid-tier gating logic. Cleanup automation needs to key off flag type/tag, targeting only release and concluded-experiment flags for staleness, not treat age alone as the signal.
- How should an experiment flag's rollout percentages differ from a release flag's, given they both look like a simple on/off toggle in code?A release flag's percentage typically only moves in one direction, toward 100%, as confidence grows, and the split is arbitrary. An experiment flag's split is usually fixed for the experiment's duration (e.g., a steady 50/50) and the bucketing must be consistent per user so the same person always lands in the same variant, because inconsistent assignment corrupts the statistical comparison.
- What's a red-flag sign that a permission flag has been mistakenly implemented and treated like a release flag?If someone proposes cleaning it up once a feature seems stable, or if it lacks the metadata tying it to a plan or segment, that's a sign it's being managed with release-flag process. A true permission flag should be tied to billing/segment data and reviewed under product/business change control, not engineering flag-hygiene sweeps.
Like the different switches on a stage-lighting board: some are cue switches you flip once during the show and then strike the set (release), some are the emergency house-lights switch that must always work even if never used (ops/kill-switch), some are for testing two lighting setups on different nights to see which the audience likes better (experiment), and some just control which rooms backstage staff with the right badge can access (permission).
saying these in an interview costs you the question
- Treats all flags as one undifferentiated bucket with the same cleanup timeline
- Proposes deleting a kill-switch flag once its feature has been stable for a while
- Doesn't know experiment flags need consistent per-user bucketing for valid statistics
- Calls permission/entitlement flags 'technical debt' to clean up
- Can't name at least two distinct flag categories unprompted