Your cold log archive rehydrates in nine hours and a live intrusion case needs eight-month-old sign-in logs — what do you do?
answer
- lag is recoverable, absence is not
- start it now, keep working
- restore the narrowest slice that answers
- which decision does it actually gate
- scoping data should never be cold
basics
~20 sStart the narrowest possible restore immediately and keep working in parallel — nine hours only costs you the case if it is serialised. Then decide whether any containment action today actually depends on the restored data.
solid answer
~50 sFirst, start the restore straight away, scoped as narrowly as possible — one source, a bounded date range, the identifiers I already have — because retrieval cost and capacity scale with what you ask for. Then work the case in parallel on what is hot: current sessions, the recent window, and the platform's own native audit trail, whose retention runs independently of my SIEM copy. Second, I state what the restore is for. If today's containment decisions do not depend on it, the nine hours are a background task rather than a blocker; if they do, the incident lead needs the ETA now to weigh acting on partial information. With one analyst, babysitting a restore burns most of a day, so it runs in the background by default. The design lesson is separate: scoping data should never be cold.
go deeper
Know that archived logs still exist but are not instantly searchable, and that the right first move is to request them and carry on rather than to sit and wait.
Be ready to explain why a narrow restore matters — retrieval cost, quotas and search-tier capacity all scale with what you ask for — and which sources have independent retention you can use meanwhile.
Show the judgment: name which decision the restore gates, give the incident lead a measured ETA, and preserve the rehydrated slice so you never restore it twice.
Own the split itself. Argue that scoping-grade sources stay hot for the whole horizon and cold storage buys depth only, and be able to defend what that costs.
## The two ways a horizon fails There are two distinct failures, and confusing them produces bad decisions. **Absence**: the records were deleted and nothing will bring them back. **Lag**: the records exist but sit in an archive tier that takes hours to make searchable. Lag is recoverable and costs wall-clock time and retrieval fees; absence is terminal. A nine-hour rehydration is lag. It is annoying, not fatal, and it becomes fatal only when someone treats it as a stop condition. ## Immediate moves **Start the restore now, narrow.** Retrieval is priced and rate-limited by volume, and a wide request can exceed the working capacity of the search tier as well as the budget. Ask for the smallest slice that answers the question: one source (sign-in logs), a bounded date range around the period of interest, and where the archive format allows it, a filter on the identifiers you already have. Ordering the whole of last spring because it is easier to type is how a nine-hour restore becomes a two-day one. **Do not serialise.** Nine hours of elapsed time costs nothing if the analyst is working; it costs a day if the analyst is waiting. In parallel: examine current sessions and tokens, the recent hot window, and any source with an independent horizon. Identity providers and SaaS platforms keep their own audit trails on their own retention schedules — often reachable through the platform's console or API when your ingested copy has aged out. **Say what the restore gates.** The right question is not `do we want the data` but `which decision changes when it arrives`. Containment of a live intrusion usually turns on present state — which sessions are active, which credentials work, which hosts are talking — not on the eighth month of history. If nothing today depends on it, the restore is background enrichment and the incident lead should be told exactly that. If something does depend on it, they need the ETA immediately so they can weigh acting on partial information against nine hours of adversary freedom. **Preserve while you are there.** Once the slice is rehydrated, hold it out of the normal lifecycle for the duration of the case. Restoring the same window twice because it re-aged is avoidable waste. ## The design lesson The scenario is a symptom of a split made on volume alone. The better split is made on question order: - **Hot for the full horizon:** low-volume, high-value sources that answer the first questions of any case — authentication and sign-in events, DNS queries, cloud and SaaS control-plane audit records, detection and alert metadata. These are small per record, so a long horizon on them is cheap relative to the bulk. - **Cold, or short:** the verbose bulk — full endpoint process telemetry, proxy records, raw network capture. These dominate the ingest bill and answer depth questions, not scoping questions. The test to apply when designing it: *can I scope an intrusion — which identities, which hosts, roughly when — without restoring anything?* If not, the split is wrong, whatever it costs. Cold storage should buy detail, never the ability to start. ## Two claims to be careful about An untested archive is a retention claim, not a capability. Restores fail on expired credentials, changed schemas, formats no longer parseable by the current pipeline, and quotas nobody has exercised in a year. Prove it periodically with a timed retrieval of a known slice, and record the actual elapsed time — that measured number, not the vendor's, is what you quote to an incident lead. And the nine hours are not free of analyst attention on a one-person team. Someone has to place the request, monitor it, re-scope it if it fails, and integrate the result. Budgeting a restore as zero-effort background work is how a small team loses a day it thought it had.
- How should the hot and cold split have been designed so this did not happen?Split on question order rather than volume. Keep the low-volume, high-value sources — authentication, DNS, control-plane audit, alert metadata — hot for the entire horizon, since they are what you need to scope any case. Push the verbose bulk, full process telemetry and proxy records, into cold or a shorter window. The design test is whether you can identify the affected identities, hosts and rough timeframe without restoring anything. Cold storage should buy depth, never the ability to begin.
- The incident lead asks whether to isolate a host now or wait nine hours for the archive. How do you frame the answer?By naming what the restored data would change. If it would only enrich the historical narrative, it is not a reason to leave a live intrusion running — say so plainly and let them act. If it would tell you whether other accounts were involved, and isolating now would tip the adversary before you know the full set, that is a genuine trade and their call to make. My job is to state the ETA, my confidence in it, and the decision it does or does not affect.
- What can go wrong with a restore you have never tested?Expired credentials or roles for the archive, quota and rate limits nobody has exercised, formats the current pipeline can no longer parse, schema drift so the fields you filter on no longer exist, and volumes that exceed the search tier's capacity. Any of these turns nine hours into days, mid-case. Run a timed retrieval of a known slice periodically and quote the measured elapsed time, not the vendor's figure.
A nine-hour restore is a slow oven. You put the dish in first and prepare everything else while it heats; standing in front of the door does not make it faster and costs you the rest of the meal.
saying these in an interview costs you the question
- Waits for the restore instead of working hot sources in parallel
- Requests a whole month or source rather than a narrow slice
- Confuses archive lag with data that no longer exists
- Treats a restore as zero effort for a one-analyst team
- Quotes an untested restore time to the incident lead as fact