A broker's audit trail is switched on. What kinds of record does it hold, and which question can it usually not answer?
answer
- two families, not one
- connections first, then decisions
- refusals are rare, and therefore cheap
- who was turned away, not who read
basics
~20 sA broker's audit trail holds connection records — who connected, from where, on which endpoint — and authorization-decision records. Most deployments keep only the refusals, so the trail rarely shows who successfully read a stream.
solid answer
~40 sThere are two families of record. A **connection record** is written when a client opens a connection: the principal the broker derived from the credential, the source address, the separately-configured endpoint it arrived on, and whether the identity was proved. An **authorization-decision record** is written when the broker checks the grant table for one attempt: principal, operation, the named stream, allowed or refused. How much of the second family exists varies — some clusters can record every decision, most are configured to record refusals only, and a managed tier emits whatever subset its operator chose. The practical consequence is worth memorising: the trail you inherit usually documents who was turned away, and not who read the data — which is the opposite of what anyone asks it afterwards.
go deeper
Remember the two families and the asymmetry: connection records say who was here, decision records say what was allowed or refused, and the usual configuration keeps only the refusals.
Be able to explain why the allowed half is expensive — the decision is taken per request on some designs and once per subscription on others, so a complete trail varies enormously in size.
Show you check what a cluster emits before designing anything on top of it, and that you know the grant table today is not the grant table that was in force during the window you are investigating.
Frame it as an evidence question for the estate: decide which streams warrant positive access evidence at all, and accept that a control depending on records a tier never produces is an assumption, not a control.
## What a broker records, and in what shape An audit trail on a broker is not one undifferentiated stream of text. Clusters that keep one keep it in two families of record, written at two different moments, answering two different questions. A **connection record** is written when a client opens a connection and presents a credential. It typically carries the principal the broker derived from that credential, the network address the connection came from, which separately-configured endpoint it arrived on, which mechanism proved the identity, and whether that proof succeeded or failed. This is the *who was here* half of the trail. An **authorization-decision record** is written when the broker evaluates the grant table for a particular attempt. It typically carries the principal, the operation attempted, the named stream the operation targeted, and the verdict — allowed or refused. This is the *what were they permitted to do* half. | Record | Written when | Answers | Silent about | |---|---|---|---| | connection record | a client connects and an identity is proved | who was here, from where, on which endpoint, by which mechanism | everything the connection went on to do | | authorization-decision record | the broker checks a grant for one attempt | which operation on which stream was allowed or refused | attempts that never reach a check, and whatever the configuration excludes | Neither family is universal in the same shape. Some brokers emit a dedicated decision record with all four fields; others fold the refusal into a generic connection-level line and never mention the resource; a rented cluster emits the categories its operator chose to publish. Treat the two families as the model, and read what your own cluster actually produces before depending on a field. ## Why the trail usually shows refusals and not reads How often a broker takes an authorization decision differs sharply between designs, and that difference decides what is affordable to record. Where each fetch or publish from a client is a separate request, the decision is taken per request, and a complete allowed-decision trail would be roughly as large as the traffic it describes. Where a subscription is established once and then served, the decision may be taken once per subscription, and a complete trail is almost free. Because the expensive case is the common one in high-throughput deployments, the ordinary configuration is one of three: record refusals only, record nothing, or record a narrow allowed subset for a few named streams. Refusals are affordable precisely because they are rare — on a healthy cluster almost nothing is refused. That economy is what produces the single fact worth carrying out of this topic: **the trail you inherit usually documents who was turned away, and not who read the data.** ## What an investigation asks, in order When someone needs to establish what happened to a sensitive stream, the questions come in this order: 1. Which identities connected to this cluster in the window, and from where? 2. Which of them were permitted to touch the stream in question? 3. Which of them actually read it, and how much? A refusal-only trail answers the first from connection records. It answers the second not from the trail at all but from the grant table as it stands today — which is not necessarily the grant table that was in force at the time. And it answers the third not at all. The gap between question three and what the trail holds is the entire reason this subject gets asked in interviews. ## The rented case On a managed tier the choice is narrower in both directions. The operator decides which categories of record exist, which fields they carry, and where they can be delivered. You can usually switch the feature on and choose a destination; you cannot add a category the operator does not emit, or populate a field it leaves empty. That makes "what does this tier actually emit" a design input rather than an implementation detail — a control that depends on evidence the tier never produces is not a control, it is an assumption. ## Four things the trail is not - **Not the stream's retained history.** The audit trail is the record of decisions about access; the stream's retained history is the data those decisions were about. They have separate lifetimes, separate destinations and separate owners. - **Not a record of activity.** A connection record proves a client authenticated. It says nothing about what that connection then did, or whether it did anything at all. - **Not proof of handling.** Even a complete allowed-decision trail ends at the broker boundary: it records that a principal was permitted to read, never what it did with what it read. - **Not evidence of absence.** No refusal for a principal is equally consistent with "it reads this stream all day" and "it was decommissioned a year ago". Only a positive record distinguishes them.
- What does a connection record tell you that an authorization-decision record does not?Where the client came from and how it proved itself: the source address, the separately-configured endpoint it arrived on, and which mechanism established the identity. A decision record starts from a principal that already exists and only says what that principal was allowed to do, so it cannot place the client on the network or explain how it got in.
- On a rented cluster, how do you find out what the audit trail will actually contain?By switching it on in a non-production environment and reading real output, not by reading the feature description. The operator fixes the categories and fields; the only reliable way to learn which of them are populated for your tier and configuration is to generate a connection, an allowed operation and a refusal, then look at what came out.
saying these in an interview costs you the question
- Assuming the trail already shows who read each stream
- Treating a connection record as proof of what the client then did
- Believing every broker records allowed decisions by default
- Expecting to add a missing field to a rented cluster's audit output
- Confusing the audit trail with the stream's own retained history