The order of volatility puts archived logs last — when is following that order wrong?
answer
- the list models one kind of decay
- an adversary is not natural decay
- value times probability of loss
- retention edges and rescheduled containers
- ordering only binds serial work
basics
~20 sWhen something other than natural decay is destroying evidence. The ranking assumes only ordinary ageing; an intruder truncating logs, a retention edge, or a container about to be rescheduled makes a durable source the one most at risk.
solid answer
~50 sThe published ranking is a decay model with one assumption baked in: nothing is actively removing evidence. The real rule is expected loss, evidentiary value multiplied by the chance the artefact is gone before you reach it, and the ranking is a good default only because those usually agree. They stop agreeing when an intruder with root is truncating the audit trail while you work, when the window you need is hours from rotating out of retention, or when the workload lives in a container the orchestrator may destroy. Then I move the at-risk source up, preserve those records, and continue down the volatile list. Two things also relax the ordering: it binds only work one responder must serialise, so independent captures run in parallel, and the cheap items take seconds, so the socket table with owning processes and the neighbour cache get taken regardless before anything long-running starts.
go deeper
Know that the ranking is a default which assumes nothing is deliberately deleting evidence, and that a person on the host breaks that assumption.
Explain the underlying rule the ranking approximates, value weighted by the chance of loss, and name two situations where a low-volatility source is actually the one at risk.
Demonstrate that you resequence under pressure with a stated reason, keep the seconds-long volatile captures regardless, and exploit parallelism rather than serialising everything by habit.
Own the standing position: which sources are forwarded off-host so they leave the intruder's reach, and therefore which decisions your responders never have to make at three in the morning.
## The assumption inside the ranking The order of volatility describes how fast artefacts decay **on their own**. Registers change constantly, memory changes with execution, the neighbour cache ages out in minutes, disk persists until retention or deletion, archives persist for months. Every one of those statements is about a system running normally. An intrusion is precisely the case where the system is not running normally. There is a person on the host who does not want the record to exist, and who may have root. The moment that is true, the decay model no longer predicts what will still be there in ten minutes, and blindly following the list means spending your first minutes on the rungs that were never in danger while the rung under active attack is destroyed. ## The rule the ranking approximates Sequence by expected loss: > **collect next whatever maximises (value of the artefact) x (probability it is gone before I get to it)** The published order is what that formula produces when the second term is driven purely by natural decay. Keep the formula and you can reason about the exceptions instead of memorising them. ## Cases where the two disagree **An intruder actively destroying.** Truncated audit records, a cleared shell history, a service log rewritten. Disk sits low on the list because disk is durable, but durability is a property of the medium, not a promise about the adversary. If you have any signal that on-disk records are being removed, preserving a copy of them becomes urgent in a way the ranking does not predict. **A short or rolling retention.** If the window you care about is close to the edge of retention on a source, that source is decaying on a clock even though the ranking calls it stable. The fix is usually a preservation request rather than a collection, which costs you almost nothing in-line. **Ephemeral infrastructure.** A workload in a container has a writable layer that vanishes when the container is replaced, and orchestration may replace it at any moment for reasons entirely unrelated to your investigation. That layer sits at the temporary-filesystem rung of the ranking, but its practical lifetime is set by a scheduler you do not control, which can make it shorter-lived than memory. **A process about to exit.** If the implant is a short-lived child, or the connection you care about is being torn down as you watch, the socket-to-process mapping has a lifetime of seconds. That is inside the ranking rather than an exception to it, but it is the case where the ranking's advice is at its sharpest: take the two-second capture before the ten-minute one. ## What relaxes the ordering rather than inverting it **Parallelism.** The ordering exists because one responder collects serially. With two responders, or with a collection that runs unattended while you work on something else, the constraint weakens: things that do not interfere can be taken at the same time, and only the genuinely serial path needs ordering. **Cost asymmetry.** Reading the socket table with owning processes, the neighbour and routing tables, the loaded module list and the process tree takes seconds in total. There is no scenario where reordering those against each other matters. The ordering question only becomes real between the cheap volatile set and the expensive long-running captures, which is where the ranking earns its keep. **Capture is not free of consequence.** Collection on a live host is observable in principle, so among equally ranked items the cheap and least conspicuous go first. That is a tie-breaker on ordering, not a reason to change the ranking. ## How to answer this in an interview Say the ranking is a default, not a law, and name its assumption out loud: it models natural decay and nothing else. Then give one concrete inversion, with the reason. The strong version is something like: on a server where the audit trail was already being truncated, I preserved a copy of the on-disk records first, then went back and took the volatile set, because the socket table was decaying at a rate set by physics while the log was decaying at a rate set by a person. The weak version is reciting the seven rungs and stopping there. Interviewers ask this question specifically to find out whether the candidate can identify the assumption under a rule they have memorised.
- You suspect the audit records are being truncated. Does that justify skipping volatile capture?No. It reorders, it does not replace. The cheap volatile set costs seconds, so you take it and then preserve the at-risk records, or run both if a second responder is available. Skipping volatile capture trades a source that cannot be recovered later for one that at least might still exist elsewhere, such as forwarded copies off the host.
- How does a second responder change the ordering?It weakens it. The ranking exists because one person collects serially; two people can take independent captures concurrently, so only work that genuinely contends needs sequencing. What does not change is that anyone acting on the host should stay clear of the volatile state another responder has not taken yet.
- If the records are forwarded to a central collector, does that make on-disk logs low priority again?Largely, and that is the point of forwarding: it moves the record out of the intruder's reach and down the volatility order legitimately. Verify the forwarding was actually working for that host and that period before you rely on it, because a source that quietly stopped reporting looks identical to a quiet host.
saying these in an interview costs you the question
- Recites the ranking as an inviolable law with no assumptions
- Ignores that an intruder with root can delete durable evidence
- Skips volatile capture entirely to save an at-risk log
- Assumes only one responder can ever collect at a time
- Treats a container's writable layer as durable because it is a filesystem