Your production flag service holds around 300 flags and nobody can say which are still needed. As the engineer setting policy, how would you run flag hygiene as an ongoing operational review, and which flags would you deliberately never expire?
answer
- stale flag equals untested code path
- owner, type, expiry at creation
- overdue count is the tracked metric
- zero evaluations means dead code
- permanent levers pay a drill tax
basics
~20 sTreat stale flags as untested code paths with a tracked count: require an owner and an expiry at creation, review overdue flags on a fixed cadence, and delete both the flag and the losing branch. Permanent operational levers are exempt but must be drilled.
solid answer
~50 sMake it a number the team watches, not a cleanup sprint. Every flag gets an owner, a type and an expiry at creation; the count of flags past expiry becomes a tracked reliability metric reviewed on the same cadence as pager load. Automate the evidence: report flags pinned at 0 or 100 percent for 90 days and flags nobody evaluates, since an unevaluated flag is dead code that outlived its call site. Removal means deleting the losing branch as well as the flag record, or you have only hidden the conditional. The deliberate exemption is operational levers — kill switches, degradation and load-shedding toggles, permission gates — which are meant to live for the life of the service. They pay a different tax: both states stay tested and the switch gets drilled on a cadence, so I cap how many the team can realistically keep honest.
go deeper
Know that a flag left in the code after its rollout finished is a branch nobody tests, and that removing it means deleting the losing code path too, not just the flag record.
Describe the mechanics that surface stale flags: required owner and expiry metadata, evaluation counts that reveal dead flags, and long periods pinned at 0 or 100 percent.
Show how the review actually runs in production — which cadence, which forum, who is named on each removal — and why an unowned cleanup item never gets done.
Commit to a policy with an explicit exempt class for permanent operational levers, bound that set by the drill capacity the team really has, and defend removal as reliability work with a real cost.
## Why 300 unknown flags is a reliability problem A stale flag is not merely clutter. It is a live conditional in production code whose disabled branch has had no traffic and no tests for an unknown length of time, and whose value can be changed by anyone with access to the flag console. Three concrete risks follow: - **Accidental exercise.** Somebody bulk-edits, restores a backup of flag state, or a default changes during a platform migration, and a branch nobody has run in two years is suddenly serving users. - **False confidence.** Responders see a flag whose name suggests it disables a risky path and reach for it during an incident. It does nothing, or something unexpected, and the outage lengthens. - **Change amplification.** Every flag multiplies the states the code can be in, so refactoring, reasoning and reproducing bugs all get harder, in proportion to a number nobody is managing. The practice question is how to keep that number honest without turning it into an annual cleanup crusade that never happens twice. ## Make it a number, then review the number The policy that works is the one that produces a metric a team already looks at. Concretely: **Metadata is mandatory at creation.** Refuse to create a flag without an owning team, a type, and an expiry date. Type is what decides everything downstream, so it has to be chosen deliberately rather than inferred later. **The tracked metric is flags past expiry.** Not total flags — total is a business fact, and a healthy team mid-rollout has plenty. Overdue flags is the number that means something has been forgotten, and it should be reviewed on a fixed cadence in a forum the team already attends, next to pager load and open action items. Reviewing it there rather than in a dedicated meeting is what makes it survive past the first quarter. **Automate the candidate list.** Two signals find most of the dead weight without human effort: flags pinned at 0 or 100 percent for a long window (90 days is a reasonable default) and flags whose evaluation count is zero. The second is the strongest evidence available — a flag nobody evaluates has outlived its call site, and its record is pure noise. **Removal means deleting the branch.** Deleting the flag record while leaving `if (true)` in the code, or leaving a hard-coded constant, converts one problem into a subtler one. The removal change deletes the conditional, the losing branch and its tests, and it is a normal reviewed code change that ships through the pipeline. ## What you never expire The policy has to have an explicit exempt class or people will quietly ignore it. Flags that legitimately live for the service's lifetime: - **Kill switches and circuit-style levers** for expensive or fragile dependencies. - **Degradation and load-shedding toggles** that turn off non-essential work under pressure. - **Permission, entitlement and licensing gates**, which are product logic expressed as configuration and are never going to be "finished". - **Regional or tenant-scoped operational controls** used during failover or maintenance. These are exempt from expiry, not from discipline. Being permanent, they attract the opposite obligation: both states stay covered by tests, and the switch is drilled on a cadence so its wiring and its fallback stay verified. That is a real recurring cost per lever, which is why the useful constraint is a **budget** rather than a blanket permission: keep the permanent set small enough that the team can genuinely keep every one of them tested and drilled. A team that can drill six levers a year should not be carrying forty and calling them controls. ## Making the default direction removal Several cheap mechanisms bias the system toward cleanup instead of accumulation: - **Ownership follows the on-call rotation**, so an unowned flag has a home rather than becoming everyone's problem. - **Expiry is visible where engineers already are** — a comment or annotation at the call site, or a build warning past the date — rather than only in a console nobody opens. - **Removal is somebody's named work item**, because unowned cleanup does not happen. This mirrors the postmortem lesson: an action item without a name is a wish. - **A cap on new flags per team**, if things are truly out of hand, forces a removal before a creation. Blunt, but it works when the number is 300 and rising. What does not work: a one-off cleanup with no ongoing metric, a bot that auto-deletes flags without a human deciding which branch wins, or a policy with no exempt class, which pushes teams to lie about flag types. ## The judgment an interviewer is testing This is a policy question with a real tradeoff, so a strong answer commits to positions and pays for them. It distinguishes temporary flags from permanent levers instead of treating all flags alike. It puts a number on the thing it wants managed, and puts that number where it will be seen. It accepts the recurring drill cost for the permanent set and bounds that set accordingly. And it is honest that removal is engineering work with a real price — you are spending time to reduce the number of states production can be in, which is a reliability investment, not tidiness.
- What single signal would you automate first against 300 unknown flags?Evaluation counts. A flag with zero evaluations over a month has no live call site, so its record can go immediately with almost no analysis. That usually clears a large fraction of the backlog cheaply and leaves you with a smaller set that needs an actual human decision about which branch wins.
- Why not just have a bot delete every flag past its expiry date?Because deleting the record is the easy half. Someone has to decide which branch survives and remove the other, and that is a reviewed code change. An auto-deleting bot leaves conditionals reading a missing flag, which resolve to whatever the default is — turning a managed stale flag into an unmanaged silent behaviour change.
- How do you stop teams from mislabelling release flags as permanent levers to dodge expiry?Make the exempt class expensive rather than free. A permanent lever owes tested off and on paths and a drill on cadence, and it shows up on the drill schedule with the owner's name. When claiming permanence adds recurring work with visible ownership, the incentive to mislabel disappears without any enforcement mechanism.
- Is there a point where the right answer is to stop using flags for a service?Yes — when the branch cost outweighs the control it buys. A service that deploys in two minutes with reliable automated rollback already has a fast undo, so flagging every change buys little and adds permanent conditionals. Reserve flags there for genuinely risky or long-running changes and for the operational levers you intend to keep.
saying these in an interview costs you the question
- Run an annual cleanup sprint and move on
- Delete the flag record and leave the conditional
- All flags must eventually be removed, no exceptions
- Total flag count is the metric to drive down
- A permanent kill switch needs no ongoing testing