A site's syslog volume fell to zero eleven days ago and nobody noticed — tampering or a benign change?
answer
- unexplained, not calm
- pin the cut to the minute
- simultaneous across device types means shared path
- does a change record match the exact timestamp
- eleven days is a blind period, say so
basics
~20 sTreat it as unexplained rather than calm. Pin the cut to the minute, scope what stopped, demand a change record whose timestamp matches, corroborate from a surface off that path. Those eleven days are a blind period.
solid answer
~50 sI would open it as an investigation with an engineering track alongside, not as a ticket for the network team. First, pin the boundary: the last event indexed from that site, to the minute. A clean simultaneous cut across every device points at something on the shared path — an egress rule, a route, a collector certificate; a staggered decay points at hosts genuinely leaving. Second, scope it: whole site, one device class, or one channel? Third, look for the innocent twin — a firewall change, a decommissioned VLAN, a site closure — but require a change record whose implemented time matches the cut and a config diff that actually blocks that port, not a plausible story. Fourth, corroborate from a surface off that path: border flow records, the collector's drop counters, directory sign-ins proving the site stayed alive. And record the blind period: no claim about those eleven days can rest on absent records.
code
text · 6 linessite d-14 d-13 d-12 d-11 d-10 ... d-1 d-0
BRISTOL 418k 402k 431k 409k 0 ... 0 0
LEEDS 377k 381k 366k 392k 374k ... 380k 371k
CARDIFF 129k 131k 127k 130k 128k ... 126k 132k
last event indexed from BRISTOL: d-11 14:07:52 (17 devices, all within 3s)go deeper
Know that a flat-zero line for a log source is a finding to raise, not a quiet period, and that you should note the last event's timestamp before anyone changes anything.
Explain how the shape of the cut discriminates causes — simultaneous across device types versus staggered — and which independent signals show whether the site was still alive.
Run it as a dual-track investigation: evidence preserved and boundary pinned before remediation, a change record corroborated to the minute, and the blind period stated explicitly in the outcome.
Own the consequence: a site can disappear from your telemetry for eleven days and nobody is accountable. Decide who is woken when a class of sources goes silent and what the SOC owes the business for the unobserved window.
## Why 'calm' is the wrong default The graph shows zero. Zero has two families of explanation, and the ordinary organisational instinct — a change broke it, raise a ticket — quietly picks one before any evidence is in. The defensible starting position is that the cause is **unknown**, because cutting the path to the collector is a cheap, quiet and entirely realistic way to remove a site from your view, and it looks identical to a misconfigured firewall rule. ## Pin the boundary Get the last indexed event from every source at that site and look at the distribution of those timestamps. - **Simultaneous, to the second, across heterogeneous devices** — switches, servers, appliances that share nothing but a network path — means something on that shared path stopped delivery. Devices do not coordinate; a path does. - **Staggered over days** means hosts individually stopped: retirement, a migration, a rebuild programme. - **One channel gone while another continues** means the failure is at generation, parsing or filtering rather than on the wire. Record the boundary before you touch anything, because the first remediation attempt will overwrite the state that answers this. ## Scope it Is it the whole site, one VLAN, one device class, or one collector? A cut that maps exactly onto a network boundary is a strong path signal. A cut that maps onto a device *type* across several networks is a strong software or parser signal. A cut that maps onto neither is the most interesting case, and usually means someone selected the sources rather than a config accidentally doing so. ## Hunt the innocent twin, then make it prove itself The benign explanations are real and common: an egress or ACL change on the path to the collector, an expired certificate on the collector's listener, a decommissioned VLAN, a closed or merged site, a re-addressing project, a collector migration that left one site pointing at a dead address. The discipline is that a plausible cause is not a confirmed one. To close on a change, you want: - an implemented timestamp that matches the last event to the minute, not merely to the day; - the config diff itself, showing a rule that genuinely blocks that destination port from that source range; - a reversal test: the volume returns when the change is backed out or an exception is added. A same-day ticket is a coincidence until it explains the exact cut. ## Corroborate from a surface off the path The decisive move is to ask a source the same actor would also have had to silence: - **Border flow records** on the path to the collector: do the syslog sessions from that site simply stop at the same second, and does anything else from that range change shape at the same time? Flow records carry the five-tuple, byte and packet counts and timestamps — they can show that sessions stopped and can never show what the payloads were. - **The collector's own counters**: connections refused, TLS handshake failures, records received and dropped. 'Nothing arrived' and 'arrived and was discarded' are different findings. - **Independent estate signals**: directory sign-ins, DHCP leases, endpoint management check-ins, badge or VPN activity from that site. If those continue through the eleven days, the site is populated and working — which kills the 'site went away' hypothesis outright. - **The change record and its author**, and whether the account that made the change behaves normally elsewhere. ## Name the blind period Whatever the verdict, the outcome includes a statement you must write down and repeat to anyone who asks: for those eleven days at that site you hold no telemetry, so you cannot assert that nothing happened there. You can assert only that you cannot see. If any partial record survives elsewhere — a copy already forwarded before the cut, a second sink, a device's own local view — that becomes the evidence for the period, and the boundary of what it covers has to be stated too. That is also the reason not to rush the restore. If tampering is still live, restoring the path tells whoever cut it that you noticed. Restoring under a narrow change, or standing up an out-of-band collection path so the original configuration stays as it was, keeps both options open. ## The finding underneath the finding Eleven days is the real defect. A site went dark and the SOC learned about it accidentally, which means there was no per-source expected-volume check with a floor, and no owner woken when a whole class of sources went quiet. Whatever the cause turns out to be, the corrective action is a silence check that would have fired on day one and an owner who receives it — otherwise the next one is also found by accident, and next time it may not be benign. ## What an interviewer is listening for That you refuse the default reading of zero; that you use the *shape* of the cut as evidence; that you demand timestamp-level corroboration rather than a plausible story; that you reach for an independent surface; and that you state the blind period instead of quietly closing the gap.
- What distinguishes a broken network path from hosts genuinely going away?Shape and simultaneity. A path change cuts every device behind it at the same second regardless of vendor or type; a decommission decays over days as hosts are retired. Then test it against independent estate signals: if directory sign-ins, DHCP leases and endpoint check-ins from that site continue through the gap, the site is alive and the path is broken, whatever the asset system says.
- The network team produces a change ticket dated that day. Is the case closed?Not on the date alone. Match the implemented time to the last event to the minute, read the config diff and confirm it actually blocks the collector port from that source range, and confirm the volume returns when it is reverted or excepted. A same-day ticket is a coincidence until it explains the exact cut and its reversal restores the flow — and it never removes the blind period from your report.
- You are ready to restore collection. Should you announce it?If tampering is still a live hypothesis, restoring the path tells whoever cut it that you noticed. Decide with the incident lead: restore under a narrow change with small circulation, or stand up an out-of-band collection path so the original configuration is preserved as evidence. Hold wider communication until you know whether you are recovering an outage or resuming an investigation.
saying these in an interview costs you the question
- Reads zero events as a quiet site
- Closes it on a same-day change ticket without matching timestamps
- Restores collection before recording the cut's boundary
- Claims the site was clean over the gap
- Checks only the collector and no independent surface