A hunt hit lands in a self-hosted CI runner's job logs. What do you preserve before you start pivoting?
answer
- two clocks: the data's and yours
- the next job overwrites the workspace
- run history is deletable by repo admins
- record the query, not a description
- subtract yourself from the timeline
basics
~20 sExport the matched records with their source, event time and collection time, and capture the query provenance: exact text, source searched, time range and timezone. Build-platform run history expires and the next job wipes the workspace.
solid answer
~50 sTwo things, and both decay on a normal business schedule rather than because anyone is hiding evidence. First the records: export the matched raw events out of the search platform to a case store, with source identifier, event time, collection time and a hash of the export — build-platform run history has a retention window and can be deleted outright by anyone with admin on the repository or organisation. Second the provenance of my own hunt: the exact query text, the index or source it ran against, the time range and its timezone, and the result count, because re-running the same query next week returns a different answer as data ages out. I also note what *I* touched, so the timeline does not later attribute my activity to the intruder. Acquisition on the runner host itself is the forensics team's job and follows its own capture order.
go deeper
Know that hunt evidence must be exported and written down, not left in a search UI, and that the query text, time range and timezone are part of the record you owe the next person.
Explain concretely why build-system evidence decays: workspaces reused by the next job, run-history retention windows, and deletion rights held by repository administrators rather than by security.
Show judgment about ordering — freeze the perishable platform records first, note your own footprints so they can be subtracted from the timeline, and leave host acquisition and pipeline stoppage to the people who own those decisions.
Be ready to argue for the retention and access you actually need on build infrastructure before an incident, given that security holds read access and the platform is owned by engineering.
## Why this step exists at all A hunter's instinct on a hit is to pivot immediately — same user, same host, wider window, upstream and downstream. Pivoting is right, but it costs you the state you were in, and on build infrastructure the underlying records are on a clock. The discipline is to spend two or three minutes freezing what you have before the pivot, because some of it will not be there when you come back. ## What is actually ephemeral here This is not a generic "logs expire" point. Build systems are unusually hostile to slow investigators: - **The runner workspace is reused.** A self-hosted runner's job directory is cleaned or overwritten by the next job on that runner. Files an intruder staged there — an archive, a dumped environment — can be gone within minutes of the next queued build, with nobody acting maliciously. - **Run history has a retention window**, often measured in weeks, after which the job log and its step output are deleted by the platform on schedule. - **Run history is deletable by non-admins of yours.** Anyone with administrative rights on the repository or organisation can delete a workflow run. They do not need to be the intruder; a tidy engineer will do. - **Shell history is a file on a host you do not own.** In this environment the security team has read access to build infrastructure but no authority to stop or freeze a pipeline, so you cannot assume you will get a second look at the host on your own timetable. - **Secret values are masked in log output**, which matters for interpretation: masking hides the value in the text, it does not mean the process was denied the value. ## What to preserve, concretely **The evidence.** Export the matched records themselves — not a screenshot — into whatever case store your team uses. Each exported record should keep its source identifier (which platform, which repository, which run and job id), its **event time** and, separately, the **collection time** the platform recorded it, because the gap between the two is how you later explain a record that looks out of order. Hash the export so you can show later that the copy you are arguing from is the copy you took. Screenshots are a fine supplement for a report; they are not the record. **The provenance of the hunt.** A finding nobody can re-derive is a finding somebody will eventually doubt. Write down, at the moment of the hit: - the exact query text, verbatim, not a description of it; - the platform, index or source it ran against; - the time range searched **and its timezone** — build systems typically log UTC while people talk in local time, and this is where investigations acquire off-by-hours errors; - the number of results, and whether the result set was truncated by a display or row limit; - the hypothesis you were testing when you ran it. Re-running the same query a week later against a rolling window is not reproduction: the window has moved and the oldest data has aged out. **Your own footprints.** Record every query you ran, every console page you opened, and above all any interactive access you made to the runner itself. Anything you do on that host lands in the same shell history and the same job environment you are going to present as evidence. An investigator who cannot subtract themselves from the timeline hands the other side an easy argument. ## What is deliberately not in scope here Two neighbouring decisions are not yours to make at this moment, and conflating them is a common answer-level mistake: - **Host acquisition.** Whether and how the runner host itself is imaged, and in what order its volatile state is captured, belongs to the forensic process and its own capture ordering. Your job is the platform-side records that will disappear on a business schedule while that decision is being made. - **Stopping the activity.** Freezing the runner takes the release train down, and in this environment you do not have the authority to do it anyway. That is an escalation to the people who own the pipeline, not a step you take between two queries. ## The failure this prevents The realistic bad outcome is not evidence tampering. It is a hunter who pivots for forty minutes, builds a compelling picture, goes to write it up, and finds the originating job run has been deleted by an engineer clearing failed builds, the workspace overwritten by the nightly job, and their own query lost in a browser tab that has since been reloaded. The finding survives only as an assertion — which, for a hit with no rule and no vendor verdict behind it, is exactly the thing it could not afford to be.
- Why record collection time separately from event time on each exported record?They answer different questions. Event time is when the platform says the thing happened, which depends on the emitting host's clock and can be wrong or manipulated. Collection time is when your pipeline received it, which bounds when the record could have been fabricated and explains apparent out-of-order arrivals. A timeline built on one of them alone will eventually be challenged on the other.
- The masked secret in the job log shows only asterisks. What can you still conclude?That the job had the secret injected and that a step read the environment containing it. Masking is a log-output filter: it redacts the text the platform prints, it does not stop the process from receiving the value. So the exposure question is settled by what ran, not by whether you can see the characters, and the credential should be treated as read.
- You want the runner's shell history but you have read access only. What now?Ask the platform owner, in writing, with the specific paths and the reason, and preserve the request itself as part of the case. Do not log in and rummage: it adds your own entries to the artefact you are asking for and it may be outside your authorisation. Meanwhile keep collecting the platform-side records you can already reach.
saying these in an interview costs you the question
- Takes screenshots and treats them as the evidence record
- Assumes job logs persist indefinitely because the platform is internal
- Describes the query in prose instead of recording it verbatim
- Ignores timezone, mixing UTC log times with local narrative times
- Logs into the runner and runs commands without recording having done so