You turn on a browser challenge for all traffic at the front door during a flood — who breaks, and how should you have sequenced it?
answer
- not every client is a browser
- callbacks and partner integrations first
- observe-only before enforce
- exempt the path, not the whole door
- owner and expiry at the moment of adding
basics
~20 sEverything that is not an interactive browser breaks: partner server-to-server calls, the payments provider's callback, embedded and no-JavaScript clients, and old mobile builds. Sequence it as observe-only first, exempt the known machine paths before enforcing, then enforce narrowly and keep a way back.
solid answer
~60 sA challenge that assumes a browser silently redefines your client population. The casualties arrive in a predictable order: the partner's server-to-server integration, which follows no redirects and keeps no token; the payments provider's inbound callback, which is a machine posting to a fixed path and cannot be asked to solve anything; an old mobile build shipping an embedded view that cannot complete the challenge; and users on assistive or text-mode clients for whom an interstitial is a dead end. The sequence that avoids most of this is: run the challenge in observe-only mode and count what *would* have failed, broken down by path and client; explicitly exempt the machine paths you find, scoped to the exact path and method rather than the whole front door; then enforce, narrowly — on the paths under attack, or only above a traffic threshold. Give every exemption an owner and an expiry when you add it, because the ones added at three in the morning are the ones still there in a year.
go deeper
Be ready to name concrete clients that cannot pass a browser challenge — a webhook callback, a server-to-server integration, an old embedded client — and why they were never going to.
Explain the rollout order and what each stage is for: observe-only to discover the machine traffic, scoped exemptions before enforcing, narrow enforcement, and a revert that needs no deploy.
Show that you would own the consequences — the support queue, the exemptions granted under pressure, and the measurement that distinguishes a working control from a control that is simply being paid for.
Be able to explain why the exemption list, not the compute, is the lasting cost of this control, and how you keep it from becoming an unmaintainable set of permanent holes.
## The shape of enforcement day The flood is live, the challenge is written and tested, and somebody flips it on for all traffic at once. Within minutes two things happen. The flood keeps arriving, because the part of it made of real hosts pays the toll. And a support queue starts filling with people and machines the challenge locked out — which is the part nobody rehearsed. The person who flipped the switch owns both. ## Who breaks, in the order they call **Machine clients that never had a browser.** A partner's server-to-server integration issues an HTTP request, reads the response, and stops. It does not follow a redirect chain, does not persist a token, and certainly does not execute a challenge. It gets the interstitial back as a body, fails to parse it, and either errors loudly or — worse — retries in a loop, adding load to the thing you are defending. **Inbound callbacks.** A payments provider posting a result to a fixed path is a machine you cannot instrument and cannot ask to change. Challenging it does not degrade a user experience; it drops financial state on the floor, and the failure surfaces later as reconciliation work. **Frozen clients.** An old mobile build with an embedded view, shipped years ago and still in use, cannot be fixed by anything you deploy today. Whatever it does when challenged is what it will keep doing. **Clients that cannot complete an interstitial.** No-JavaScript browsers, text-mode clients, and assistive setups where a redirecting, self-refreshing interstitial breaks the flow. These users rarely file tickets; they simply leave, which is why they do not appear in the incident channel and do appear in the conversion numbers. ## The sequence that would have avoided most of it 1. **Observe-only first.** Run the challenge logic and record the verdict without acting on it. What you want out of this is a breakdown of what *would* have been challenged, by path and by client shape — and specifically a list of paths where the traffic is entirely non-browser. That list is your exemption list, discovered rather than remembered. 2. **Exempt the machine paths before enforcing, and scope them tightly.** Exempt the specific callback path and method, not the whole front door; require that path to carry its own credential, so the exemption is not simply an open door. An exemption expressed as a source-address allowlist across everything is the widest possible hole, and it is exactly the hole an attacker looks for once the challenge is public. 3. **Enforce narrowly.** On the paths actually under pressure, or only above a traffic threshold, so normal days look normal and the blast radius on enforcement day is bounded. 4. **Keep the way back.** One switch, known to the on-call, that returns the front door to its previous behaviour without a deploy. ## What the list becomes The exemption list is the durable cost of this control, not the CPU. Entries get added under outage pressure, by whoever had access, with no note about why. Six months later there are dozens, nobody can name what each protects, and removing one risks breaking something invisible — so nobody does. The only cheap moment to prevent that is the moment of adding: every entry gets a named owner and an expiry date at creation, and periodically the whole set is re-run in observe-only mode to see which entries would still fail today. Entries that would pass are deleted. ## Measuring whether it worked Watch three things, and be careful about the direction of each claim. The rate at which challenges are issued tells you how much traffic is unproven, not how much is hostile. The solve rate tells you the population can pass, not that it is friendly. The one that matters is completion of real user journeys — sign-up, checkout — because that is the only number that reflects the people the control locked out. A high solve rate with unchanged upstream volume is the signature outcome: the challenge is working exactly as designed, the flood is paying the toll, and the problem has become a capacity problem.
- You exempt the partner integration by source address. What does the flood do with that?It becomes a free path — anything arriving from that range skips the challenge, and the range is discoverable by anyone who probes. Scope the exemption to the exact paths and methods that integration uses and require its own credential there, so the exemption's blast radius is smaller than the whole front door.
- Six months on, the exemption list has forty entries. What went wrong and what do you do?Entries were added under outage pressure with no owner and no expiry, so nobody can now say what a removal would break. Fix it forward: attach an owner and an expiry to every entry, then re-run the challenge in observe-only mode against current traffic and delete every entry that would pass today.
- Challenges are being solved at a high rate and upstream volume has not moved. What is your read?The control is working as designed and the flood is made of real hosts paying the toll. That is not a tuning failure and a harder challenge will not fix it; the remaining levers are capacity, upstream absorption, and narrowing what the flood is permitted to reach.
- Which metric tells you the challenge is hurting real users?Completion of genuine user journeys — sign-up and checkout conversion — compared against the period before enforcement. Issued and solved counts describe the traffic that engaged with the challenge; the people it turned away are visible only as journeys that stopped happening.
saying these in an interview costs you the question
- Enables the challenge everywhere at once with no observe-only pass
- Assumes every legitimate client is an interactive browser
- Exempts a partner by source address across the whole front door
- Adds exemptions during the outage with no owner or expiry
- Calls it a success because the solve rate is high
- Has no way to revert without a deploy