skip to content

Using the management-API audit record, how do you reconstruct the timeline of an unexplained overnight change to a shared service?

level: seniorimportance: must knowfreq 55%

answer

  1. bound the window first
  2. target plus one hop out
  3. denials and describes are evidence
  4. join calls by correlation id
  5. a gap is a lead, not an alibi

basics

~20 s

Bound the window between the last known-good observation and detection, filter to calls targeting that resource and its dependencies, keep denied and read-only calls too, join them by correlation id, and allow for delivery lag.

solid answer

~50 s

Work outward in four steps. **Bound the window** from the last moment the configuration was observed correct to the moment it was noticed wrong — the record is only as useful as the interval you search. **Filter by target**, taking the resource itself plus the things that can change it indirectly: the access grants that govern it, the automation that deploys it, the settings above it. **Keep everything, not just the successes**: denied calls show attempts, and describe-style read calls immediately before a change often reveal someone or something orienting itself. **Join the calls up** by correlation id, then by principal plus source address, so a multi-step operation reads as one action rather than five. Two cautions: delivery into the searchable record is not instantaneous, and a change made outside the management API produces no entry at all — a gap is a lead, not an alibi.

go deeper

for a junior

Know that the platform keeps a searchable record of management API calls, and that finding a change in it means searching a time window and filtering by the resource that changed.

for a middle

Explain the mechanics: bounding the window from a known-good anchor, filtering by target, and why denied and read-only calls belong in the result set rather than being filtered away.

for a senior

Demonstrate the judgment — widening by one hop to grants and deploying automation, joining calls into actions, and stating out loud that delivery lag and out-of-band paths make a gap a lead rather than an alibi.

for a principal

The angle you own is what makes this repeatable: whether the window you can search is longer than the time it typically takes anyone to notice, and who is expected to run this rather than improvise it at 08:00.

## What you are actually building The deliverable is not "who did it". It is an **ordered list of calls**, each with a principal, a time, an outcome and a target, that accounts for the state you are looking at now. Naming a culprit is something a human does afterwards, from that list plus context the platform never had. Keeping that distinction explicit is what separates a useful timeline from an accusation, and it is also what stops you stopping too early — the first plausible entry you find is rarely the whole sequence. ## The four moves 1. **Bound the window.** Find the last observation that proves the configuration was still correct — a successful deploy, a monitoring sample, a screenshot, a previous audit query — and pair it with the moment someone noticed it wrong. Everything else is search inside that interval. A window that is too narrow is the most common reason an investigation concludes "nothing in the record", and widening it is cheap. 2. **Filter by target, then widen by one hop.** Start with calls whose target is the resource itself. Then add the things that can change it without touching it: the grants that govern who may call it, the automation or pipeline identity that deploys it, and any setting applied at a level above it. A configuration that "changed itself" is nearly always something one hop away being reapplied. 3. **Keep the calls that did nothing.** Two categories get filtered out by reflex and should not be. **Denied** calls record attempts the access model stopped, and a cluster of them is often the shape of the story. **Describe-style read calls** — listing, getting, showing — are how both a person and an automation orient themselves before acting, so a burst of them immediately before the change usually tells you which principal was working in that area. 4. **Join the calls into actions.** One human action is frequently several API calls. Use the correlation id where the record carries one, then fall back to principal plus source address plus a short time bucket. Presented un-joined, five entries look like five events and invite the wrong conclusion. ## The three traps | trap | what it looks like | what to do | |---|---|---| | delivery lag | the record shows nothing for the last few minutes | wait and re-query before concluding; delivery into the searchable store is not instantaneous | | out-of-band change | no entry exists anywhere in a generous window | ask whether the change even travelled through the management API — a setting inside the service's own interface leaves no management event | | ordering within a second | several entries share a timestamp | do not infer causality from adjacency at the same resolution; use the correlation id or the semantics of the calls | To those add a smaller one: timestamps are normally recorded in a single reference timezone while your incident is discussed in local time. Converting once, at the start, prevents an hour of arguing about a change that happened when everyone was asleep. ## What the timeline hands off A finished timeline is the input to three other conversations, and trying to answer them from the record itself is where investigations go wrong: - **Was that caller permitted to do it, and should the grant have been that wide?** The record names the principal and says the call succeeded. Whether the permission model should have allowed it is the access-model question, and it is answered by reading grants, not events. - **Does the live resource now differ from what was declared?** That comparison belongs to whatever declares the resource, and it answers a different question: not who called, but what no longer matches. - **What do we change so it cannot happen again?** That is the postmortem, and it needs the human context the platform never saw. ## A worked shape A shared reporting service starts returning short history at 08:00. The last successful deploy at 17:40 the previous day is the known-good anchor, so the window is fourteen hours. Filtering to that resource yields one configuration-changing call at 02:41 by a deployment automation's identity. Widening by one hop shows, six minutes earlier, a series of describe calls from an interactive session on a different address, and a denied attempt to change the same setting directly. Joining by correlation id shows the 02:41 call was one of three from a single pipeline run. The timeline that comes out is: a person looked, tried to change it directly, was refused by the access model, and then ran the pipeline — which carried the change through with credentials that were allowed. Notice what the record established and what it did not. It established the sequence, the identities and the outcomes. It did not establish that the person intended to bypass the refusal, and a timeline that asserts that has stopped being evidence.

  • You find no entry for the change anywhere in a generous window. What do you conclude?
    Not that nothing happened. Three hypotheses, in order: the change did not travel through the management API — a setting inside the service's own interface, or a system layered above the platform, produces no management event; the query window or the target filter missed it; or delivery had not completed when you looked. Rule them out in that order before concluding anything.
  • Why should describe-style read calls stay in the timeline when they changed nothing?
    Because they show orientation. Both a person and an automation typically list and inspect before acting, so a burst of read calls from one principal just before a change is strong evidence of who was working in that area — and if the change itself came through a shared automation identity, those reads may be the only entries that point at a specific session.
  • The change arrived under an automation's identity. How does that alter the shape of the investigation?
    It moves it off the platform. The record ends at the credential; who triggered the automation is in that system — a pipeline run, a schedule, a merged change. So the timeline hands over at that boundary, and the next artefact you ask for is the automation's own run history, correlated by the same timestamp.

saying these in an interview costs you the question

  • Searches a window that starts at the moment of detection
  • Filters to successful calls and drops denials
  • Treats an absent entry as proof no change was made
  • Reads five API calls as five separate human actions
  • Names a person from a shared automation's identity
  • Concludes intent from the order of two adjacent entries