skip to content

An unopened cloud detection is 94 days old and your control-plane logs retain 90. What now?

level: seniorimportance: should knowfreq 44%

answer

  1. the row survived, the evidence did not
  2. undetermined is a real disposition
  3. history is gone, the present is not
  4. retention is not uniform across sources
  5. measure age against the shortest horizon

basics

~20 s

The item is permanently unanswerable: the telemetry needed for a verdict expired while it waited. Close it as undetermined, never as benign, and pivot to what survives — current resource state and any longer-retention source.

solid answer

~50 s

Separate the alert from its evidence. The detection record still exists in the case system, which is why the row looks workable; the control-plane events around it are gone, so the questions that decide the verdict — who called it, from where, what else that identity did, what happened next — can no longer be asked. The correct disposition is *unable to determine, evidence expired on day 90*, and writing anything that reads as **no evidence of malicious activity** is a claim you cannot support. Then pivot to what has not expired. Events expire; **state does not**. If the detection said a disk snapshot's create-volume permission was granted to an account id outside the organisation, go and look now: is that permission still there, is the account id one you recognise, does the snapshot still exist, is the identity still active. Longer-retention sources — billing and cost records, configuration inventory, an organisation-level trail in a separate account — may also outlive the main index.

code

json · 12 lines
json
{
  "case_id": "CLD-2026-018842",
  "detection": "snapshot create-volume permission granted to an account id outside the organisation",
  "account": "prod-data-847201",
  "severity": "medium",
  "status": "unopened",
  "raised": "2026-03-04T02:11:09Z",
  "first_opened": "2026-06-06T09:40:00Z",
  "age_days": 94,
  "control_plane_log_retention_days": 90
  ...
}

go deeper

for a junior

Be ready to recognise that an alert record outliving its supporting logs means no verdict is reachable, and that the honest disposition is undetermined rather than anything that sounds like a clean result.

for a middle

Expect to explain why the evidence and the alert have different lifetimes, and to name sources whose retention may differ — organisation-level trails, configuration inventory, billing records, source and change history.

for a senior

Demonstrate the pivot from expired history to current state: query whether the shared permission, the external account id and the identity still exist today, and be precise that a present-state finding does not establish what happened in March.

for a principal

Own the consequence for reporting. Decide that permanently unanswerable items are counted and shown to the business as lost knowledge rather than absorbed into a queue-length number that implies everything was eventually handled.

## What has actually expired Two different things carry the word *alert*. The **detection record** lives in the case-management system and can persist for years: rule name, timestamp, severity, a handful of fields the rule copied out. The **evidence** — the control-plane trail, the sign-in records, the flow logs, the endpoint telemetry — lives in a searchable index for a fixed window. When the item passes that window unopened, the row is still clickable and the investigation is over before it starts. That is what makes a 94-day-old item in a 90-day estate a genuinely different object from a 40-day-old one. The available outcomes have collapsed from three (malicious, benign, benign-true-positive) to one: **unknown, permanently**. ## The disposition you write is the whole test Under time pressure the tempting close is *no evidence of malicious activity found*. It is true in the trivial sense — you found no evidence — and it is deeply misleading, because the reason is that no evidence exists to find, not that you looked and the estate was clean. Anyone reading the queue later, an auditor included, will read it as a verdict. Write instead: **unable to determine — control-plane evidence for this detection expired at day 90; the item was never opened.** Record the date the window closed and which sources were needed. That closure is honest, it is checkable, and, unlike a fabricated benign verdict, it aggregates into a number the organisation can act on: *this many detections became permanently unanswerable last quarter.* ## Events expire; state does not The most valuable move a senior analyst makes here is the pivot from history to the present. The log of *what was done* is gone. The *result* of what was done is generally still sitting in the estate, queryable today, with no retention window at all. For a snapshot shared outward, that means going and checking now: - Does the snapshot still exist, and is its create-volume permission still granted to that external account id? - Is that account id one of yours, a known partner's, or unrecognised? Does the same id appear on any other resource, in any other account, in any of the three providers? - Does the identity that appears in the detection still exist, is it still active, and what can it do today? - Is there a change record, ticket, or infrastructure-as-code commit from that week that explains the sharing? A configuration repository or a build system keeps its own history on its own schedule and may well cover March. A present-state finding is not a verdict about the past, and you should say so plainly: seeing the permission still in place tells you the sharing exists, not who created it or whether it was approved. But it converts a dead item into a live one. If the sharing is unrecognised and unexplained, you are no longer closing an old alert — you are opening an investigation into data that may be readable by a third party right now, and the fact that you cannot reconstruct March does not reduce that. ## Also check what is retained differently Retention is rarely uniform. Organisation-level or security-account trails are often kept far longer than per-account ones. Billing and cost records, configuration-inventory snapshots, backup catalogues, ticketing systems and source repositories all have their own horizons, and in a multi-provider estate they never line up. A senior candidate lists two or three of those rather than declaring the case dead at the first expired index. ## What this changes about how you measure the backlog The operational lesson generalises beyond this one item. A queue's real deadline is not a service-level target somebody wrote down; it is the **shortest retention among the sources a verdict would need**. So the backlog metric that matters is not the raw count of unopened items but two derived numbers: the age of the oldest unopened item measured against that horizon, and how many items are within a short window of crossing it. Those are the ones you can still save this week, and they are the ones a sweep should target. And the number to publish afterwards is the count of items that crossed — detections that will never have a verdict. That is not a backlog statistic, it is a description of what the organisation cannot know. ## What a strong answer sounds like Name the collapse (evidence gone, verdict impossible), refuse the benign close, pivot to current state and differently-retained sources, and finish with the measurement change: track age against the retention horizon and report the items that crossed it, rather than reporting a queue length.

  • The snapshot is still shared with that external account id today. Is that enough to declare an incident?
    It is enough to open an investigation and to act on the resource. What it establishes is that the sharing exists now, not who created it or whether it was approved. Look for an owner, a change record, an infrastructure-as-code commit, and the account id elsewhere in the estate. If nothing explains it, treat it as data exposed to an unknown third party and work it as an incident, while stating plainly that the March activity itself can no longer be reconstructed.
  • What exactly do you write in the case when you close it?
    Unable to determine, with the reason: the item was never opened, and the control-plane evidence required to adjudicate it expired at day 90. Record which sources were needed and the date the window closed. Never a phrase like no evidence of malicious activity, which reads as a verdict when the truth is that no evidence remained to examine.
  • Which backlog numbers should the SOC track instead of the raw unopened count?
    The age of the oldest unopened item measured against the shortest retention horizon a verdict would need, the count of items within a short window of crossing it, and the count that already crossed. The first two are actionable this week; the third is the honest description of what the organisation can no longer know.

saying these in an interview costs you the question

  • Closing it as no evidence of malicious activity found
  • Assuming the surviving alert record means the evidence survived
  • Declaring the case dead without checking current resource state
  • Treating retention as uniform across sources and providers
  • Reporting queue length while items silently cross the horizon

context