skip to content

Why does a management-API audit trail record the creation of an object store but not each read of an object inside it?

level: middleimportance: should knowfreq 44%

answer

  1. two halves, two kinds of event
  2. shape of the estate against use of it
  3. volume ratio decides the default
  4. data events opt-in and scoped
  5. you cannot enable it retroactively

basics

~20 s

Creating the store is a control-plane call and is recorded by default; reading an object is data-plane traffic, which is orders of magnitude more frequent and is usually recorded only if you switch that recording on in advance, per resource.

solid answer

~50 s

The platform splits its record along the same line as the platform itself. **Management events** describe calls that change or describe the shape of the estate — create, update, delete, list — and providers record them by default because the volume is modest and the value is high. **Data events** describe use of the thing once it exists: an object read, a key used to decrypt, a row fetched. Those run at a rate many orders of magnitude higher, so providers generally leave them off, scope them to chosen resources or paths, charge for them separately, and often deliver them to a different destination. The consequence is retrospective and painful: after a suspected leak you can see who created the store, who opened its access policy and when, but not who read the bytes — unless somebody turned data recording on before the incident.

go deeper

for a junior

Recall that the platform records calls that create and change resources by default, and that recording every read of stored contents is a separate, optional thing.

for a middle

Explain the split along the control-plane and data-plane seam, and why volume — thousands of management calls against millions of reads — is what puts the default where it is.

for a senior

Demonstrate the consequence: after a suspected leak you can produce the configuration history and the exposure window, but not the list of objects read, unless someone enabled that recording beforehand.

for a principal

The decision you own is where contents-level recording is worth its cost across the estate, and writing that classification down so an incident starts by checking a list rather than by hoping.

## The split follows the platform's own split A cloud platform has two halves. The **control plane** is the management API: the surface through which resources are created, reconfigured, described and destroyed. The **data plane** is the thing you rented actually doing its job: serving objects, answering queries, decrypting, routing packets. The audit record is split along the same seam, and knowing which half an action fell on tells you immediately whether an entry is likely to exist. - **Management events** are calls against the control plane. Creating an object store, changing its access policy, enabling versioning, deleting it. These are recorded by default across the account. - **Data events** are operations against the contents. Reading one object, writing one, listing the contents of a prefix, using a key to decrypt a payload. These are typically *not* recorded unless you ask for them. ## Why the default goes that way It is a volume argument before it is anything else. An account might see thousands of management calls a day; the same account's object store might serve millions of reads an hour. Recording every one of those produces a stream that is larger than the workload's own traffic in event count, and the platform would be charging you to store a near-duplicate of your application's access pattern. So providers converge on the same design, with differences in the details: | dimension | management events | data events | |---|---|---| | on by default | yes | no, in general | | typical volume | thousands per day | millions per hour | | scoping | whole account | per resource, often per path or prefix | | charged | usually bundled | usually charged per event recorded | | destination | the account's audit record | often a separate destination you choose | Providers differ on how far the scoping goes — some let you select a subset of operations as well as a subset of resources, and some record a small set of especially sensitive data operations by default — so the safe statement in an interview is that the *default is off or narrow* and the enabling is a deliberate act, not that any particular operation is or is not covered. ## The consequence you must be able to state This is a decision you can only make **before** the incident. The shape of the trap is always the same: 1. A store is created, and its management events are recorded automatically. 2. Months later somebody changes its access policy to something permissive. That change is a management call, so it *is* in the record, with the caller and the time. 3. Data is suspected to have been taken during the window between that change and its discovery. 4. You go looking for who read what — and there is nothing, because data events were never switched on for that store. You can therefore answer *how the exposure became possible, who caused it and exactly when* with complete precision, and you cannot answer *what was actually taken*. That asymmetry is worth saying out loud in an interview, because it is what drives the decision: you enable data recording selectively, on the stores whose contents would matter in that conversation, and you accept that you cannot afford it everywhere. ## How to choose where to turn it on Since blanket coverage is not affordable, the choice is a classification exercise rather than a technical one: - **Turn it on where the contents are the asset** — the store holding customer records, the key material, the export destination of a reporting pipeline. - **Scope it as tightly as the platform allows** — a path or a prefix rather than the whole store, and read operations rather than every operation, where reads are the risk. - **Leave it off for the bulk** — build artefacts, static assets, scratch space and anything whose disclosure costs nothing. - **Write down the decision** so that the next incident starts by checking the list rather than by hoping. ## The adjacent confusion Two other records look like this one and are not. A service's own access log — the request log the thing you rented emits about traffic it served — is a feature of that service, in its own format, delivered where you point it. And your application's logs describe what your code did with the data after it had it. Both can be valuable, and neither is the platform's audit trail: they are written by something inside the boundary, they can be turned off by whoever operates the workload, and their format and retention are yours to get wrong. Pipelines, parsing and retention design for all of those streams belong to the logging discipline rather than to the platform's own record.

  • After a suspected leak from an object store with only management events recorded, what can you still establish?
    The complete history of the store's configuration: who created it, who changed its access policy and exactly when, who enabled or disabled versioning, and every denied attempt against its management surface. That gives you the exposure window and the responsible call. What you cannot produce is the list of objects actually read during that window.
  • Why is recording data events for every store a bad default even if budget were not a problem?
    Because the value per event collapses while the cost of finding anything rises. A record in which almost every entry is a routine read is one you cannot search quickly during an incident, and the retention you can afford for it is shorter. Selective recording on the resources that would matter keeps both the signal and the search usable.

saying these in an interview costs you the question

  • Assumes every read of stored data is recorded automatically
  • Thinks data event recording can be enabled retroactively for a past window
  • Confuses the service's own access log with the platform's audit record
  • Says the split exists for security reasons rather than volume
  • Believes management events must also be switched on before they appear