skip to content

You must upgrade the one inline IPS at each of forty branches -- how do you sequence the bypass windows an intruder could ride?

level: seniorimportance: should knowfreq 37%

answer

  1. never blind everywhere at once
  2. pilot on the cheapest site
  3. measure the window, do not estimate it
  4. a permanent 02:00 slot is a readable schedule
  5. no local hands means a van

basics

~20 s

Never all at once and never on a permanently fixed slot. Pilot on the least valuable branch to learn the real window, stagger the rest, put the most sensitive sites last, and cost the callout before you start.

solid answer

~50 s

Refuse the obvious plan: one maintenance hour for all forty leaves the whole estate uninspected at once and puts an unproven image on forty boxes nobody can reach. Pilot on one low-value branch and measure the window end to end -- inspection stopping to inspection provably resuming -- rather than trusting a stated boot time. Then stagger, so only a slice of the estate is blind at a time, and vary the slot between cycles so the window is not a standing appointment anyone inside can plan around. Order by what each site carries: the branch you least want unjudged goes last, shortest window, most people awake. Per branch, confirm which posture the relay is wired for and that its owner accepted it. And with no local hands, agree beforehand what happens if a box does not come back.

go deeper

for a junior

Know that an inline box's upgrade is an outage or a blind spot depending on how the relay is wired, and that upgrades at unattended sites are scheduled and announced rather than done ad hoc.

for a middle

Explain why the window is measured from inspection stopping to inspection provably resuming, and why staggering sites changes the size of the blind area even though each site's own window is unchanged.

for a senior

Demonstrate a real rollout order: pilot, stagger, sensitive sites last, upstream telemetry watched live, positive verification afterwards, and a costed answer for a box that does not return.

for a principal

Own the estate-level trade: how much of the estate may be unjudged at once, who funds the spare units and the callouts, and what the change record must show for an auditor asking about coverage gaps.

## The plan to reject first The instinct with forty sites is one maintenance window: announce 01:00 to 03:00, push everywhere, go to bed. It is the wrong shape for three reasons. 1. **It makes forty small holes into one estate-wide hole.** Whatever your inspection normally buys you, you have none of it anywhere for that hour. A staggered plan means the blind area is a slice, not the map. 2. **It removes the pilot.** A bad image, a configuration that does not survive reboot, or a relay that latches is now a problem at forty unattended sites at once, and every fix needs a vehicle. 3. **It concentrates the human attention you have.** Two people cannot watch forty windows. They can watch four. ## Measure the window before you schedule it The number that matters is not the vendor's boot time. It is the interval from the moment inspection stops to the moment inspection is demonstrably back: power down or service stop, firmware write, reboot, rule set compiled and loaded, interfaces back in path, and a positive check that traffic is being judged again. On the first branch, time all of it and record it. That measured number drives everything else -- how many branches fit in a night, what you tell branch owners, and whether the window is short enough that a fail-closed posture is even discussable. ## Ordering Order by blast radius, not by convenience or alphabetically: - **First:** a branch whose traffic is low value and whose downtime is cheap. This is the pilot; treat it as a test, not as work completed. - **Middle:** the bulk, staggered a few per night, with each night's results feeding the next. - **Last:** the site you least want running unjudged, or whose downtime is most expensive. By then the image is proven, the real window length is known, and you can afford to give that site the shortest, best-attended window. ## Do not let the window become a schedule anyone can read A window that is always the first Sunday at 02:00, forever, is a published fact inside the organisation and to anyone who has been inside it long enough to notice. That does not require a sophisticated adversary: it only requires someone who has watched a few cycles. Vary the slot between cycles, keep the detailed per-branch order in fewer hands than the change announcement, and accept that this is a mild control rather than a strong one -- the primary defence is a short window, not a secret one. The more realistic exposure is duller and more certain: sessions already open ride straight through, because nothing evaluates traffic while the relay is closed, and an engine that comes back to a connection it never saw start has very little to say about it. ## Posture, per branch, confirmed before the night For each branch, know which way the relay is wired and confirm the site's owner has accepted it. Fail-open means the branch trades through the window with no inspection; fail-closed means the branch cannot trade until the box returns. It is a legitimate answer for these to differ across the estate, and it is not a decision to discover at 02:10 while looking at a dark site. ## Watch from upstream, because the box will not be reporting During each window the engine reports nothing, so your visibility is whatever sits around it: router or WAN flow records (five-tuple, packet and byte counts, timestamps, no payload), DNS resolver logs, proxy logs where the branch is proxied. Have those in view for the window rather than after it, both to notice the branch has actually gone quiet on schedule and to be able to reconstruct what crossed. ## Plan the failure you cannot fix remotely The specific risk at an unattended site is that the box does not come back or the relay does not restore. Decide in advance: does someone drive tonight, does a courier take a spare in the morning, or does the branch stay as it is until Monday? Also decide what 'as it is' means -- a latched-closed relay leaves the branch working and uninspected, which is the failure that produces no complaints and therefore no urgency. Every window should end with a positive verification that the box is in path and judging traffic, not with the absence of a phone call. ## What good sounds like in an interview Staggered, piloted, ordered by what each site carries, measured rather than estimated, watched from upstream while the box is down, verified positively afterwards, and with the truck roll costed before anyone touches a box.

  • Why not run all forty windows in a single maintenance hour and be finished?
    It turns forty small gaps into one estate-wide gap with no inspection anywhere, and it removes the pilot: a bad image or a latched relay lands simultaneously at forty sites nobody can reach. Staggering also lets the first branch teach you the real window length before you commit the rest.
  • What do you watch during a window, given the box itself is reporting nothing?
    Everything upstream of it: router or WAN flow records, DNS resolver logs and the proxy if the branch is proxied. They give addresses, ports, byte counts, names asked for and timestamps, and no payload -- enough to say what crossed and when, never what it was. Watching live also confirms the window started and ended when you think it did.
  • How do you choose which branch goes last?
    By blast radius rather than convenience. The site whose traffic you least want unjudged, or whose downtime costs most, goes after the image has been proven elsewhere, gets the shortest measured window, and gets the most people awake and watching. The pilot goes first for the mirror-image reason.

saying these in an interview costs you the question

  • Schedules the entire estate in one maintenance window
  • Uses the same fixed slot indefinitely for years
  • Treats the stated boot time as the real window
  • Has no plan for a box that does not come back
  • Ends the window on the absence of complaints rather than a check

context