On-call widened a host's egress firewall rule at 03:00 to end an incident and the baseline now fails: what did that cost the control?
answer
- the host is cheap, the period is not
- a pass is a point sample
- gap bounded by two measurements
- incident timeline narrows the window
- a reboot can erase the evidence
basics
~20 sThe host is easy to fix; the period is not. The control can no longer be claimed continuously effective, and the gap is bounded only by the two measurements unless something independently timestamped the change.
solid answer
~50 sSeparate the host from the claim. The machine's state is knowable and correctable, and the on-call engineer did the right thing at 03:00 — ending the outage beat preserving a firewall ruleset. What broke is the assertion the control supports: it held at the last passing measurement and does not hold now, so the interval between them contains an unknown-length gap, and the control is compliant-with-a-gap rather than continuously effective for the period. Narrowing the gap means finding something that timestamps the change independently: the incident timeline, the modification time on the ruleset, an agent's convergence log. There is a second cost worth naming — if the rule was widened only in the running firewall and never saved, a reboot restores the compliant ruleset and erases the only evidence, so the machine passes again while the period still contains the gap.
go deeper
Know that fixing the firewall rule is only part of the answer, and that the failure also says something about the time between the two measurements.
Be ready to explain why the gap is bounded by the measurements rather than by the incident, and how a runtime-only firewall change differs from a saved one.
Demonstrate that you can state both bounds honestly, name where independent timestamps come from, and recognise a self-closing finding as evidence lost rather than a problem solved.
Own the framing that a correct incident decision became a finding because nothing recorded it. The remedy is a cheap capture path or moving the control off the host, not discipline for the engineer.
## The cost is not the host The instinct is to treat this as a host problem: the egress rule is too wide, narrow it, close the finding. That part is real but cheap. The expensive damage is to what the control lets you say. A preventive or detective control is only interesting as a statement about a **period**: this egress restriction was in force on this host for the quarter. Every measurement is a point sample, and a re-failure converts the statement from *held throughout* to *held at the last passing sample, does not hold now, unknown in between*. That interval is the gap, and its length is not a property of the incident — it is a property of how far apart the samples are. The host is a one-line fix; the period cannot be repaired retroactively by any amount of remediation. ## Bounding the gap honestly The defensible statement has two bounds and you should be explicit about both: - The **outer bound** is what the measurements alone support: non-compliant at some point within the interval between the last pass and this failure. - The **inner bound** is what independent evidence supports: if the incident timeline puts the firewall change at 03:12 on a specific night, the gap starts then, not at the previous measurement. A senior answer reaches for the second bound and names where the evidence would come from — the incident record, the modification time on the saved ruleset, the shell history on the host, the converging agent's log, the change ticket if one exists. If none of those exist, say so plainly rather than quietly asserting the narrower window. Reporting *at most eleven days, at least since the incident at 03:12* is a stronger position than either an unbounded shrug or a confident number you cannot support. ## The evidence can delete itself Here is the detail that separates people who have operated hosts from people who have not. Firewall state is layered: the ruleset live in the kernel and a saved ruleset restored at boot. If the on-call engineer widened only the running ruleset — which is exactly what happens when you are trying to unblock traffic in seconds — the change is in force but never persisted. The next reboot restores the compliant saved ruleset, the machine passes the following measurement, and the only trace of the gap is gone. The finding closes itself and the period still contains the hole. Audit configuration behaves the same way: rules loaded into the kernel at runtime disappear on reboot, and rules written to disk do nothing until they are loaded. So *it passes now* is not evidence the gap did not happen, and a control that only ever compares against current state cannot see its own history. This is the argument for keeping the observed values from each measurement, not just the verdicts. ## Do not blame the engineer The organisational read matters as much as the technical one. The engineer widened the rule because the alternative was a longer outage, and any process that makes that the wrong call has mispriced the risk. Decay of this kind is a consequence of legitimate work. The defect is not that a rule was widened under pressure — it is that the system had no cheap path for the change to be *recorded* as it happened, so a correct operational decision silently became a compliance finding discovered days later by a machine. That reframing changes what you ask for afterwards. Not *stop touching production*, but: can the change be captured at the moment it is made, can the measurement interval be short enough that the outer bound is not embarrassing, and is this control one that should live somewhere an incident cannot reach it at all. ## The second-order effect on the claim One more consequence is worth stating because interviewers listen for it: a re-failure resets any continuity clock. If the control's value came from *continuously in force since March*, that phrase is no longer available and the earliest date you can restart it from is the first measurement that passes after the gap closes. That is why the same finding lands very differently on a host nobody claims anything about and on a host inside a scope where the claim is the product. ## What a strong answer covers Host fixable, period not; the outer and inner bounds and where independent timestamps come from; the running-versus-saved split that can erase the evidence; and the refusal to treat a correct incident decision as negligence.
- The host passes the next measurement without anyone touching it. What do you conclude?Not that nothing happened. A runtime-only firewall change vanishes at reboot, restoring the saved compliant ruleset, so a self-closing finding is a strong hint that the change was never persisted. The gap in the period still occurred; the evidence for it is simply gone. Rely on the incident timeline and the retained observed values rather than the current verdict.
- How would you narrow the gap from the outer bound to something defensible?Find an independent timestamp for the change: the incident record, the modification time on the saved ruleset, shell history on the host, or a convergence log entry showing when the agent last enforced the resource. State both bounds explicitly — at most the interval between measurements, at least from the timestamped event — and say when no independent evidence exists.
- The host owner argues the on-call engineer should be written up. How do you respond?Push back. Ending the outage was the right call, and a process that punishes it will get slower incident response and hidden changes rather than fewer findings. The defect is that there was no cheap way to record the change as it was made. Aim the follow-up at capturing the change and at whether this control belongs somewhere an incident cannot reach.
saying these in an interview costs you the question
- Treats fixing the host as the whole remediation
- Claims the gap started at the failing measurement
- Says a passing re-run proves nothing was wrong
- Blames the on-call engineer for ending the outage
- Ignores that runtime changes vanish on reboot