skip to content

A /16 you pushed from a threat feed at 08:58 took the CRM offline at 09:05 — what do you do?

level: seniorimportance: should knowfreq 47%

answer

  1. confirm it was your push, quickly
  2. smallest reversal that restores service
  3. never trade the block for blindness
  4. re-enter as a host, not a range
  5. allow-list the range with an owner

basics

~20 s

Confirm from the deny logs that your push caused it, then make the smallest reversal that restores service. Put the removed range straight into matching so it is still watched, and re-enter the indicator narrowly.

solid answer

~40 s

First, confirm rather than assume: the proxy deny records should show the CRM destination denied against my list name, starting minutes after the 08:58 push, only in the regions that received it. That takes a minute, not twenty. Then restore service with the smallest reversal that works — remove the /16 entry if I can identify it immediately, otherwise revert the push and re-add the good entries afterwards. The business is down and I am the cause, so I bias hard to restoring. Critically, I do not trade the block for blindness: as the prefix leaves enforcement, the same range goes into matching so contact is still alerted. Then I re-enter narrowly — the specific hostname I needed — allow-list the shared range with an owner, and re-run it alert-before-block.

go deeper

for a junior

Be ready to say you would check the deny logs for your own list name against the affected destination, and that restoring the business service comes before defending the block.

for a middle

Explain the reversal choice — pull the single entry versus revert the push — and why the removed range must immediately be matched for alerting so visibility is not lost with the block.

for a senior

Show the whole arc: fast causation check, minimal reversal, detection retained, narrow re-entry verified against your own traffic, allow-list entry with an owner, and alert-before-block before enforcement returns.

for a principal

Own the process change rather than the incident: prefix-shaped entries get a collateral check against internal traffic before release, and pushes to a fleet with no staging go region by region.

## Minute zero to two: is it you? A CRM outage seven minutes after a list push is your change until proven otherwise, but 'until proven otherwise' still means looking. The evidence is cheap: - Proxy or firewall **deny records** naming your list or rule object, with the CRM's destination as the target. - Timing: denials starting after the push completed, not before. - Blast pattern: only the regions that received the push, only outbound to that destination — not a general failure of the CRM for everyone including users off the corporate path. If a user on a network that never received the push is also broken, it is not your change and you have just avoided a pointless reversal. Spend a minute or two here, no more. Being confidently wrong in either direction is expensive: leaving your own block in place prolongs an outage, and reverting a change that was not the cause tells you nothing and loses the enforcement. ## Minute two to twelve: the smallest reversal that restores service Two options, and the choice is about speed of identification: 1. **Remove the single entry.** If the deny record names the matched entry and you can pull that /16 in a minute, do that. Everything else you pushed stays enforced. 2. **Revert the push.** If you cannot name the offending entry quickly, roll back to the previous list state and re-add the good entries afterwards. It costs you the rest of the batch temporarily, and under time pressure with a business system down that is the correct trade. The governing principle is that you are the cause and the service is revenue-bearing. This is not the moment to defend the block, and 'security outranks the CRM' is the answer that ends the interview. ## The step people forget: do not trade the block for blindness Removing an enforcement entry removes the *prevention*. It must not remove the *observation*. As the prefix comes out of the block list, the same range goes into matching against proxy, DNS and flow records so any internal contact still raises an alert. Otherwise the rollback hands the adversary destination back its reachability and takes away your view of it in the same action, and if the intrusion is live that is exactly the window they use. This is also what lets you answer the question that comes next: during the seven minutes the prefix was enforced, did anything internal actually get denied toward the destination you were aiming at? The deny records from those minutes are evidence and are worth keeping. ## After service is back: re-enter narrowly - Replace the prefix with the **specific** indicator you actually needed — the hostname or the full URL where you have it, a single address only if you believe the host is dedicated. A prefix was never the right shape. - **Verify** the replacement does not contain the CRM's destination, by checking your own proxy logs for what your users actually reach in that range rather than trusting the feed's description. - **Allow-list the range** with an owner and a written reason, so the same entry cannot arrive again next month and be pushed by someone who was not here today. - Run the replacement **alert-before-block** for a fixed window before enforcement is turned back on, and go live one region at a time. ## And the person you broke The sales director whose team lost the CRM at the start of a Monday is owed a plain factual account: what was pushed, what it hit, when it was removed, and what stops it recurring. Keep it factual and short — you caused it, so do not narrate it as a security success. The durable fix she should hear about is the process one: an entry containing a prefix rather than a host does not go out without a collateral check against your own destination logs first, and pushes to a fleet with no staging behind it go to one region ahead of the rest. ## The lesson underneath The outage was not caused by an intelligence failure. The feed may well have been right that adversary infrastructure lived somewhere in that range. It was caused by enforcing at a **granularity** the enforcement point could not aim, in an estate where the only place a change first runs is production. Both of those are fixable without touching the feed at all.

  • Why not simply revert the entire push and be done with it?
    It works and under pressure it is sometimes right, but it also removes every other entry in that batch, including ones that were doing something. Prefer pulling the single entry when the deny record names it and you can act within a minute or two; revert wholesale when you cannot identify it that fast. Restoring service wins either way.
  • The intrusion is still live. Does that change whether you roll back?
    It changes what replaces the block, not whether you restore the business. Pull the prefix, immediately re-enter the specific destination you actually needed, and put a detection on the whole range so any contact is still visible. Leaving a revenue system down to preserve an over-broad block is not containment, it is collateral you chose.
  • How do you stop the same prefix being pushed again in a month?
    The range goes on the allow-list with an owner and a reason, so a future push collides with a recorded decision rather than an absence. And any candidate entry that is a prefix rather than a host gets checked against your own proxy and flow logs for internal traffic before it leaves — the check that would have caught this one in seconds.
  • What evidence should you preserve from the seven minutes the block was live?
    The deny records themselves. They tell you whether any internal host was actually trying to reach the destination you were aiming at, which is a finding in its own right, and they are the timeline evidence for why the outage started and stopped when it did.

saying these in an interview costs you the question

  • Leaves the block in place because security outranks the CRM
  • Reverts the entry and leaves no detection on the range
  • Assumes the outage was not the block without reading deny logs
  • Re-adds the same prefix later with no allow-list entry
  • Spends twenty minutes proving causation while users are down

context