skip to content

A WAF stopgap outlived the vendor fix it covered — what evidence would let you delete the rule, and what does producing it cost?

level: seniorimportance: should knowfreq 46%

answer

  1. prove it at request level
  2. bypass the rule, then test
  3. release notes are not evidence
  4. log-only before deletion, with a date
  5. no archived sample, no payable proof

basics

~20 s

Only request-level evidence counts: with the rule bypassed, the fixed build must reject the archived incident sample and its variations on its own. Release notes and a quiet partner are not evidence. It costs a testable build from the vendor, an archived sample, and a window in which you carry the risk.

solid answer

~60 s

The rule's job was to stop input reaching an unfixed handler, so the only proof it is redundant is that the handler now refuses that input without it. Replay the archived incident sample, plus the variations you know of, against the fixed build with the rule bypassed — a staging instance of the exact production build, or a header-gated test path — and observe the origin's own response and side effects. A 400 or an empty result from the application, with no reflected error and nothing written, is evidence; a passing pipeline in the vendor's repository is not. Then flip the live rule from blocking to log-only for a window that spans your rarest legitimate flow, so you learn what real traffic it was still matching before you remove it. The costs are concrete: the vendor must supply a testable build, you carry the residual risk during the log-only window, and if the original request sample was never archived the burden of proof may simply be unpayable — which is itself the finding to report.

code

text · 12 lines
text
# rule 100217 (virtual patch, incident INC-4471) bypassed for this replay
# archived sample: POST /portal/search, param q, content-type form

build 4.2.1 (pre-fix, staging)
  -> 200, 11.4 kB, response body reflects a backend error string
  -> application log: query executed

build 4.6.0 (current production build, staging)
  -> 400, 312 B, {"error":"invalid search term"}
  -> application log: input rejected before query
...
# same replay repeated for the JSON body and multipart variants

go deeper

for a junior

Know that removing a stopgap needs evidence about the application, not about the rule, and that the incident's request sample is the artefact everything later depends on.

for a middle

Explain the replay: the same request against old and fixed builds, the rule bypassed, comparing the origin's own response and side effects rather than the proxy's verdict.

for a senior

Demonstrate a staged retirement — replay, then a log-only window sized to the rarest legitimate flow, then deletion with a written record — and be honest about what the evidence does not cover.

for a principal

Own the residual risk during the window and the contractual gap that made a testable build hard to get. Decide what the organisation does when the evidence is unobtainable.

## The claim you have to substantiate Deleting a stopgap rule is a claim: *the application now rejects, on its own, the input this rule was written to stop.* Everything else people offer is a proxy for that claim and none of it substantiates it. - "The release notes say it is fixed" — a statement of intent by someone who cannot see your route. - "The partner reports no problems" — evidence about legitimate traffic, not about the defect. - "The vendor's tests pass" — evidence about the vendor's tests. None of these are request-level. The rule sits at request level, so the evidence must too. ## The replay The artefact that makes this possible is the **archived request sample** from the original incident: the route, the method, the content type, the parameter, and the value. If it was kept, retirement is a half-day of work. If it was not, you are trying to prove that a fix addresses an input you can no longer state, and that is the honest answer to give. With the sample in hand: 1. Get an instance of the **exact production build** you can send requests to — a staging deployment of the same version, or a test virtual host, with the stopgap rule **bypassed** on that path. Bypassing is essential; testing through the rule tests the rule. 2. Replay the sample and the variations you know of, and record the **origin's own response**: status, size, whether an error is reflected back, and whether anything changed on the application's side. 3. Compare against the same replay on the old build, so the difference is attributable to the fix rather than to a change in the test. A clean result looks like the application rejecting the input itself — a validation error, an empty result set, no reflected diagnostic — where the old build behaved differently. ## Then stop blocking before you stop matching Even with the replay in hand, do not go from blocking to absent in one step. Flip the live rule to **log-only** and leave it for a window long enough to span your rarest legitimate flow — a monthly batch, a quarter-end export. Two things come out of that window: - what real traffic the rule was still matching, which occasionally reveals that a legitimate partner request has drifted into matching it and has been failing quietly for months; - whether the route is being probed, which tells you what your exposure was while you were relying on the rule alone. The trap is that log-only is comfortable, and a rule left in log-only forever is the same unowned artefact as before with less protection. Put a date on the window. ## What it costs, stated to whoever is asking | Cost | Who pays it | |---|---| | A testable instance of the fixed build | the vendor, usually under a contract clause that does not exist yet | | Engineering time to replay and compare | your team | | Residual risk during the log-only window | whoever owns the partner relationship, explicitly | | Keeping the evidence for the next auditor | your team, forever | That middle row is why this is a negotiation and not a task. Somebody has to accept that for two or six weeks a bug that was once exploited is guarded by nothing but a log line, and that acceptance should be written down with a name and a date on it. ## Scope your conclusion honestly The replay proves the fixed build rejects **that input on that route**. It does not prove the bug class is gone from the application — that is work for whoever owns the code, and it is not what your rule was ever claiming. Say precisely what you proved, delete the rule, and keep the replay output, the bypass configuration and the log-only window's counters with the closure note. The next person to ask "why is there no rule here?" deserves a better answer than "someone deleted it." ## The finished record The deletion is done when there is a written record tying together: the incident, the rule ID, the archived sample, the replay result on the old and new builds, the log-only window and what it saw, the named approver, and the date. That record is the thing that stops the next stopgap on this virtual host from becoming permanent, because it demonstrates that removal is possible and shows exactly what it takes.

  • The original request sample was never archived. What do you report?
    That the deletion criterion is unpayable as things stand: you cannot demonstrate the fixed build rejects an input you can no longer state. The options are to reconstruct the sample from the rule's own pattern and the incident write-up, to ask the vendor what the fix actually changed and derive a test from that, or to keep the rule and record explicitly that it is retained because the evidence was lost — which is a decision someone signs, not a default.
  • Why must the rule be bypassed rather than left in place during the replay?
    Because with the rule active you are measuring the rule, not the application. The request never reaches the handler, so a rejection tells you nothing about whether the fix works. The bypass has to be scoped tightly — a test virtual host, a header-gated path, a staging instance — so you are not opening the production route to the very input you are replaying.
  • How long should the log-only window run before you delete?
    Long enough to span the rarest legitimate flow that touches the route — if a partner submits a batch monthly, a two-week window proves nothing about it. Set the length from the traffic pattern, not from convenience, put an end date on it, and treat an expired window with no decision as a failure of the retirement rather than a safe default.
  • What does the replay actually prove about the application?
    That this build rejects this input on this route. It does not prove the bug class is absent from the rest of the application, and it does not prove the next release keeps the fix. Say exactly that when you close the ticket, and if the route matters enough, ask the vendor for a regression test on their side so a future build cannot silently reopen it.

saying these in an interview costs you the question

  • Accepts release notes or a vendor assurance as proof
  • Tests through the rule instead of bypassing it
  • Deletes straight from blocking mode with no observation window
  • Leaves the rule in log-only indefinitely and calls it retired
  • Claims the replay proves the whole bug class is gone
  • Has no archived request sample and proceeds anyway

context