An egress log shows a nightly build reached an unexpected host - how do you find which stage?
answer
- attribution has to be designed in advance
- per-job identity at the egress point
- intersect with per-step timestamps
- proxy, DNS and flow logs together
- bytes out versus bytes in
basics
~20 sAttribution needs a per-job egress identity - a source address or proxy credential that puts the job id in every log line - intersected with retained per-step timestamps. Bytes sent versus received then separates exfiltration from a downloaded payload.
solid answer
~50 sAttribution is only possible if you built for it. Give each job a distinguishable identity at the egress point - its own network namespace and source address, or a per-job proxy credential so the proxy writes the job identifier into every line - then intersect the connection timestamp with the job's retained per-step timing to land on a specific stage. Pull three sources, because none is sufficient alone: proxy access logs for hostname and byte counts, DNS query logs which often survive when the connection itself was refused, and flow logs from the job's namespace which catch non-HTTP traffic that ignored the proxy. Bytes sent versus received is the field that matters: a large upload means something left, a small request returning a payload means something arrived and you now need a copy of it. Finally, diff the resolved dependency set for that build against the previous night to find the newly introduced package.
go deeper
Know that a build talking to an unknown host is worth investigating, and that the logs which answer it live off the runner, not on it.
Explain what the three log sources each cover and why bytes sent versus received changes the conclusion you draw about the connection.
Demonstrate that attribution is a design decision made beforehand - per-job egress identity and retained step timing - and show how you get from a connection to the responsible dependency.
Own the retention and instrumentation argument: what you spend so that a future incident is scoped to one pipeline instead of an estate-wide credential rotation and rebuild.
## The two questions you have to answer An unexpected destination in the egress log for a nightly build turns into exactly two questions, and everything you did beforehand decides whether they are answerable: 1. **Which job, and which step inside it, opened that connection?** 2. **What moved - did something leave, or did something arrive?** ## Attribution: you need a per-job egress identity The usual dead end is a log line that says a runner in the CI address range talked to an unknown host at 02:14. If every job leaves through one shared address, that is where the investigation stops. What makes attribution possible is giving each job a distinguishable identity at the egress point: - a distinct network namespace or source address per job, - a per-job proxy credential or token so the proxy writes the job identifier into every line, - or an identifier injected as a request header by the proxy client the runner uses. With that, the connection resolves to one pipeline run. Getting from the run down to the **step** needs the job's own timeline - per-step start and end timestamps retained alongside the logs - so you can intersect the connection time with the step that was executing. This is the concrete reason to keep step-level timing rather than a single job duration. ## Three log sources, and why one is not enough - **Proxy access logs** give hostname, method, status, and crucially bytes sent versus bytes received. They only cover traffic that actually went through the proxy. - **DNS query logs** often survive when the connection itself was refused - the lookup happened even though the socket never opened - so they catch attempts the proxy never saw, and they are where name-encoded exfiltration shows up. - **Flow or connection-tracking logs** from the job's network namespace catch what neither of the above does: non-HTTP protocols and direct-to-address traffic that ignored the proxy setting. Ship all three off the runner in real time. An ephemeral runner is destroyed at the end of the job, and anything only on its disk is gone before anyone reads the alert. ## Reading the traffic **Bytes sent versus bytes received is the first field to look at.** A large upload to an unknown host is exfiltration; a small request that returns a few hundred kilobytes is a second-stage download, and that changes what you do next - in the second case something ran that you do not have a copy of, and recovering it from the proxy cache or from the artifact becomes the priority. ## Tying it back to a cause The other half of the answer is knowing what code was running. If you retain the **resolved dependency set for every build**, you can diff the affected night against the previous one; a package that appeared or moved version at that point is your first candidate. Without that record, you are comparing manifests that both say the same range and learning nothing. ## What the absence of logging costs Say this plainly, because it is the argument that funds the work. With no per-job attribution and no retained logs you cannot scope the incident, so you must assume the worst case: every credential the job could have reached is treated as compromised and rotated, and every artifact produced in the window from the earliest possible compromise is suspect and gets rebuilt. You also cannot answer a customer or a regulator asking what left your network. Logging turns an estate-wide rotation and rebuild into a one-pipeline problem. ## Retention Keep egress logs, DNS logs and the per-build resolved dependency record for at least as long as the artifacts those builds produced remain in service, and at minimum long enough to cover the gap between a compromise and its public disclosure - which is routinely weeks or months. The question you will be asked later is not `what is happening now` but `was the build we shipped in March affected`, and only retention answers it. ## One thing people forget **Log the denies as loudly as the allows.** A refused connection from a build job is both your best alert and a piece of evidence, and a policy that silently drops packets without recording them throws away the more useful half of default-deny.
- What does the absence of these logs actually cost you during the incident?Scope. With no attribution you cannot say which pipeline or which window was affected, so you assume the worst: rotate every credential the job could have reached and rebuild every artifact produced since the earliest possible compromise. You also cannot answer anyone asking what left the network.
- Why keep the resolved dependency set for every build, not just the manifest?Manifests usually express ranges and read identically from night to night. The resolved set records exactly which versions and digests were used, so you can diff the affected build against the previous one and see the package that appeared or moved - which is what turns an unexplained connection into a named cause.
- How long should egress and dependency-resolution logs be retained?At least as long as the artifacts those builds produced remain in service, and long enough to cover the usual gap between a compromise and its disclosure. The question you get later is whether a build shipped months ago was affected, and only retention answers it.
saying these in an interview costs you the question
- Assumes proxy logs alone identify which job connected
- Relies on logs left on an ephemeral runner that is destroyed
- Ignores DNS logs because the connection was blocked anyway
- Cannot distinguish data leaving from a payload arriving
- Logs allowed connections but silently drops denied ones