Your rented-edge WAF degrades during an attack, and failing away to the origin removes the control — how do you decide this beforehand?
answer
- the decision has a TTL attached
- someone else's outage, your status page
- you cannot un-publish an address
- a fallback nobody exercises is a hypothesis
- decide the posture while nothing is burning
basics
~20 sDecide the posture before the incident, because during it you have a DNS TTL you cannot shrink retroactively, an attacker already engaged, and an irreversible act: publishing the origin address cannot be undone. Choose in advance between staying degraded behind the edge and a fallback position you already run.
solid answer
~50 sThe reason this has to be settled in advance is that every option is worse mid-incident. Failing away means pointing the public name at the origin: you wait out the TTL you set weeks ago, you serve traffic with no judgment in front of it while an attacker is actively probing, and you permanently publish an address that passive DNS archives keep forever — you cannot un-publish it. The realistic choices are: accept degraded service and stay behind the edge; or pre-provision a position you operate that can carry the traffic with the rules still applied, paying for idle capacity and for keeping two rule sets equivalent. What I cannot do is test the provider's failure — it is someone else's failure domain, their outage lands on my status page, and I cannot fix it. So what I write down beforehand is the trigger, who is authorised to call it, the real TTL, and what the fallback does and does not judge.
go deeper
Know that a control run by somebody else is a dependency you cannot repair, and that removing it to restore service means requests stop being judged. Availability and coverage are being traded here.
Explain the mechanics: the public name resolves to the provider, the cutover is a DNS change bounded by a TTL set in advance, and pointing at the origin both bypasses judgment and reveals the address.
Demonstrate the pre-agreed posture — trigger, caller, measured cutover duration, what the fallback judges — and the habit of exercising the fallback on a calm day so it is not a hypothesis.
Own the trade at the service level: which services may fail closed and stay dark, which justify paying for a second position you operate, and what you are willing to state publicly about a dependency you cannot fix.
## What you actually bought when you rented the position A control *ahead* of you — operated by someone else, reached because your public hostname resolves to their addresses — comes with a property the two self-operated positions do not have: **you cannot test its failure and you cannot fix it**. Their incident is your incident, their status page becomes the explanation on yours, and the thing you would most like to do during the incident is exactly the thing you have arranged not to be able to do cheaply. That is not an argument against renting the position. It is the price of it, and the senior skill is having priced it before the day it is charged. ## Why the mid-incident decision is bad in three separate ways **The TTL was set weeks ago.** Redirecting the public name to the origin takes effect over the record's time-to-live, and lowering it now does not help — resolvers already hold the old value for its full lifetime. If the plan requires a fast cutover, the low TTL had to be in place before anything happened, and a permanently low TTL costs query volume and adds a dependency on resolution working during the incident. **The attacker is already engaged.** Whatever the edge was refusing is now aimed at an unjudged origin. You have removed the control at the moment its subject was most active, and you have also removed whatever record of those requests the edge was producing. **Publishing the origin is irreversible.** Once the public name resolves to the origin address, third-party DNS history keeps that association indefinitely. When the edge recovers and you point the name back, the address is still known and still answers. You did not open a door for an hour; you opened it permanently and closed the sign-post. ## The postures worth choosing between | Posture | What it costs | What it gets you | | --- | --- | --- | | Stay behind the degraded edge | Availability, and an outage you cannot shorten | Coverage never lapses; the origin stays unpublished | | Pre-provisioned position you run | Idle capacity, and two rule sets kept equivalent | A path you can switch to that still judges requests | | Fail away to the origin | Unjudged traffic plus permanent exposure of the address | Availability now, at a price you pay forever | The middle row is the one that requires work on a good day. A fallback proxy you operate has to carry equivalent rules, or the day you use it you either block things production allows or allow things production blocks — and rule drift is silent until then. It also has to be exercised, which means deliberately taking the edge out of the path on a calm morning and watching what breaks, because a fallback that has never carried real traffic is a hypothesis. ## What you write down before the incident This is the artefact that turns judgment into an operable decision: 1. **The trigger** — what evidence, from your own vantage rather than the provider's status page, means the edge is unusable. 2. **Who calls it** — a named role who can accept the exposure, not the on-call engineer improvising at three in the morning. 3. **The mechanics and their real duration** — the exact change and the TTL that governs it, measured, not assumed. 4. **What the fallback judges** — and, explicitly, what it does not, so nobody discovers the gap while it is live. 5. **The way back** — how you return, and the acknowledgment that the origin address stays known afterwards. ## What you can and cannot verify about the rented position You can verify your own side: that every path traverses the edge, that your rules behave the way you think, that the cutover works and how long it truly takes, and how your service behaves when the edge returns errors rather than going dark — degraded is more common and harder than down. You cannot verify their failure modes. What substitutes for testing is contractual and historical: published availability, incident communication commitments, and their record in past outages. Treating that substitution as equivalent to a test is the mistake; naming it as a risk you accepted, with a fallback you have exercised, is the answer an interviewer is listening for.
- What can you actually test about a rented edge before you depend on it?Only your own side: that every path really traverses it, that your rules behave as expected, that the cutover works and how long it takes in practice, and how your service behaves when the edge returns errors rather than disappearing. Their failure modes are untestable, so what replaces the test is their published availability, their incident communication commitments and their outage history — which you accept as a risk, not as evidence.
- If you pre-provision a position you run as the fallback, what does it cost?Capacity that earns nothing most of the year, and rule drift: two rule sets that must stay equivalent or the day you cut over you block what production allows, or allow what production blocks. It also has to be exercised on a calm day by deliberately taking the edge out of the path, which is a small planned risk traded against an untested one.
- Why is lowering the TTL once the incident starts not a fix?Resolvers already hold the old record for its full remaining lifetime, so a change now applies only after they expire what they cached. The short TTL had to be in place beforehand, and keeping one permanently costs query volume and makes the cutover depend on resolution behaving normally during the event.
saying these in an interview costs you the question
- Treats switching DNS to the origin as the standard fallback
- Assumes the provider's failure modes can be tested
- Thinks lowering the TTL during the incident speeds the cutover
- Believes exposing the origin address is temporary and reversible
- Keeps a fallback position that has never carried real traffic