Why run a newly published C2 domain back through five months of historic DNS logs?
answer
- intelligence arrives after the fact
- blocking only protects going forward
- the past lives in stored logs
- earliest match anchors the timeline
basics
~20 sBecause indicators arrive late. A block added today only matches traffic from today onward; the intrusion the indicator describes may already be months old in your stored logs. The retro sweep is the only way to see backwards.
solid answer
~40 sIntelligence almost always arrives after the activity it describes. Pushing a newly published command-and-control domain into the DNS and proxy blocklists protects you from tomorrow; it says nothing about yesterday. If a trojanised build-tool update ran on your CI runners in October and the vendor only published the second-stage domain in March, the only place that beacon still exists is in your retained DNS and proxy records, so a retro search is what converts "we are now protected" into "we know whether we were hit". It also anchors everything downstream: the earliest historic match tells you which hosts and which time range to examine, and roughly how long this has been going on. Skip it and an intrusion that already happened gets quietly filed as prevented.
go deeper
Be ready to say plainly why a new indicator is searched backwards as well as blocked forwards, and to name which logs you would search first for a domain name.
Explain the mechanics: which sources actually carry a domain name, how far each realistically reaches, and what the earliest match does and does not establish.
Show the judgment: sweep before hosts get rebuilt, use the earliest match to scope rather than to conclude, and recognise that on transient infrastructure the log may be the only surviving evidence.
Own the standing arrangement — which classes of intelligence automatically trigger a sweep, the time you commit to running it in, and how that commitment is funded rather than argued out per report.
## The two directions an indicator travels When a threat-intelligence report lands with a command-and-control (C2) domain in it, that single artefact is used in two completely different directions, and confusing them is the mistake this question exists to catch. **Forward**, into the controls: the domain goes into the DNS sinkhole, the proxy blocklist, the mail gateway, the detection rule. From the moment enforcement takes effect, anything in the estate that tries to reach that name is blocked or alerted on. This is *prevention and detection of future activity*. **Backward**, across the telemetry you already hold: the same string is searched against months of stored resolver logs, proxy records, flow records and host telemetry. This is the **retro search** (also called a retrospective sweep or retro-hunt), and it answers a question the forward path structurally cannot: *did this already happen to us?* ## Why intelligence always lags An indicator becomes publishable only after somebody, somewhere, has already been attacked with it, investigated it, and decided to disclose. In a supply-chain case the lag is often brutal: a signed update from a build-tool vendor ships in October and is trusted by everyone who takes it; the vendor discovers the compromise in February; the second-stage domain the implant calls is published in March. Every organisation that installed the update had five months of activity that was, at the time, invisible to every control they owned — because the artefact that would have identified it did not exist yet. That lag is not a defect to be engineered away. It is the normal shape of intelligence, and it is precisely why the backward direction exists. ## What the retro search is searching Start with the surfaces that carry the artefact and reach furthest back at the lowest cost. For a domain that is usually the resolver query logs and the egress proxy, both of which record a name and a client for every request. Flow records show sessions to addresses rather than names, which makes them useful once you know what the domain resolved to at the time. Host telemetry — process and network events from an endpoint agent — gives you the process that made the request, but typically covers a much shorter window than the network logs do. In an engineering estate the delivery vector matters too. If the vector was a poisoned dependency or a trojanised installer, the build system's own records — job logs, dependency resolution, artefact provenance, and the package registry's record of which version was pulled and when — are small, cheap and directly on point. ## What a hit does for you The **earliest** historic match is the highest-value output of the sweep. It gives you: - a **lower bound** on when the activity was already underway, which anchors the incident timeline; - a **scope**: the set of clients that produced matches is your initial host list; - a **pivot window**: whatever preceded that first match on those hosts is where the delivery event lives. Note the phrase *lower bound*. The earliest match is the earliest evidence in the telemetry you happen to still hold, not the moment of initial access. Access may well predate your first record, and a candidate who states the earliest match as "the time the attacker got in" has overstated the evidence. ## The asymmetry that makes this urgent Hosts are transient in a modern engineering estate. CI runners are rebuilt from images, containers are recycled, laptops are reimaged, cloud instances are replaced. Five months later the machine that made the request very often does not exist in any recoverable form. **The log may be the only surviving evidence of the intrusion**, which is why the sweep is run before anybody starts touching or rebuilding hosts, and why the record of what was found is treated as evidence rather than as a scratch query. ## How this differs from hunting A retro search starts from a **known artefact** handed to you by somebody else and asks whether it appears in your data. It is mechanical, bounded, and its scope is fully determined by the indicator and the window. That is a different activity from hypothesis-driven hunting, which starts with a behaviour you suspect and no artefact at all. Both look like "searching old logs"; only one of them begins with something a vendor published. ## The wrong answers "We blocked it, so we are covered" is the classic miss: blocking is a statement about the future. "We would have alerted at the time" is the same error wearing a detection badge — a rule that did not exist in October produced nothing in October. And "we searched from the day we ingested it" is a sweep that has been silently truncated to zero useful history.
- Which telemetry would you sweep first for a C2 domain, and why that order?Resolver query logs and the egress proxy first: both carry the name itself, cover the whole estate, and usually reach furthest back for the least query cost. Flow records come next, once you know what the name resolved to at the time. Host process telemetry adds the process that made the request but normally covers a shorter window. In an engineering estate I would also sweep the build system and package-registry records, because they are small and speak directly to a poisoned update.
- You have the earliest historic match. What do you do with it?Treat it as an anchor, not a verdict. It gives a lower bound on when activity was already happening, the initial host list, and a window to pivot into — what those hosts installed or ran just before it. I would widen the sweep either side of that timestamp on other surfaces and then look for the delivery event. I would not report it as the moment of initial access, because it is only the earliest thing my retained telemetry can show.
Adding the indicator to a blocklist is fitting a lock to the door. The retro search is reviewing the last five months of camera footage to see who already walked through it.
saying these in an interview costs you the question
- Says blocking the domain closes the question
- Only searches forward from when the indicator was ingested
- Treats the earliest log match as the moment of first access
- Assumes intelligence is real-time and never lags the activity
- Rebuilds affected hosts before the historic sweep has run