Recording every allowed authorization decision on a busy broker is usually rejected. What drives that cost, and what narrower configurations keep some value?
answer
- two costs, not one
- the check sits on the serving path
- volume follows the request shape
- each narrower row gives up specific evidence
basics
~20 sThe decision sits on the request-serving path, so a complete trail costs work per authorized operation and produces volume comparable to the traffic itself. Narrower options: refusals only, a named set of sensitive streams or principals, one record per session, or sampling.
solid answer
~50 sTwo costs, not one. The authorization check happens while the broker is serving the request, so writing a record adds work inside the broker's own latency budget; and where the check is taken per request rather than once per subscription, the trail approaches the size of the traffic it describes. Writing it back onto the same cluster makes it worse, because audit records are themselves writes that generate further decisions. The narrower configurations trade a specific piece of evidence for volume: **refusals only** catches probing and misconfiguration but is blind to abuse by an over-granted principal; **a named subset** of sensitive streams or principals is proportional to what you actually care about; **one record per session** tells you a principal touched a stream but not how much; **sampling** yields statistics, not evidence about a named access. On a managed tier you may get none of these choices.
go deeper
Know that a complete record of every allowed operation is usually not switched on, and that the reason is the sheer number of decisions a busy broker takes rather than the size of each record.
Explain both costs — work on the serving path and volume that tracks the request count — and name at least two narrower configurations with the evidence each one sacrifices.
Show that you would measure against the busiest hour, keep refusals unconditionally, and use a time-boxed full-recording window rather than flipping a permanent switch during an incident.
Own the choice of which streams warrant complete access evidence at all, and state the estate rule up front, since the subset has to be chosen before the incident that needs it.
## Where the cost actually is It is tempting to read "record every allowed decision" as a storage problem. It is two problems, and the first is the one that gets the proposal rejected. **The check is on the serving path.** A broker consults the grant table while it is answering a client. Turning a decision into a durable record adds work — formatting, buffering, and eventually a write to somewhere — inside the latency budget of the operation the client is waiting on. Implementations buffer and batch to keep it off the critical path, which converts the latency cost into a memory cost and a loss window: whatever is buffered when a node dies is the part of the trail you do not have. **The volume follows the request shape, not the data shape.** How often a decision is taken varies by design. Where every fetch or publish is a separate request, the decision is per request, and one audit record per authorized request means a trail with a record count comparable to the operation count of the cluster. Where a subscription is established once and then served, the decision may be taken once per subscription, and complete recording is almost free. The same feature is therefore cheap on one platform and unaffordable on another, which is why blanket advice about it is wrong as often as it is right. **The destination compounds both.** If audit records are written back into a stream on the cluster they describe, they are themselves writes — which are themselves authorized, which produces further decisions to record. That feedback is bounded in practice by excluding the audit stream from recording, but the load is not: you have added traffic to the cluster whose behaviour you were trying to observe. ## The narrower configurations, and what each one gives up | Configuration | Volume | What it proves | What it cannot show | |---|---|---|---| | Refusals only | Very low | Who was turned away, and when a client is probing or misconfigured | Any access by a principal whose grants already permit it | | Named subset of streams or principals | Proportional to the subset | Complete access evidence where it was decided it matters | Anything outside the subset, chosen before the incident | | One record per session | Low | That a principal touched a stream at all, and when | Volume, repetition, and the shape of the access | | Sampling | Tunable | Statistical patterns across principals and streams | Whether one named access happened — the evidential question | The ordering matters. Each row down the table buys volume reduction with a specific piece of evidence, and the decision is about which evidence you are willing to be without when someone asks a year later. 1. **Start from the question.** Decide which streams would trigger an investigation if they were read by the wrong principal. That list is almost always short, and it is the natural subset. 2. **Keep refusals unconditionally.** They are the cheapest row and they carry a different signal from the others — a rising refusal count is a client that was reconfigured, a credential that was replaced badly, or someone trying doors. 3. **Choose the session-level option deliberately, not by default.** It answers "did this principal ever touch this stream" and nothing finer. For a grant review that is often enough; for an investigation into how much data left, it is not. 4. **Treat sampling as measurement, not as evidence.** A sampled trail is useful for seeing which principals dominate access to a stream. It cannot exonerate anyone and it cannot convict anyone. ## The rented case, which is not a configuration at all On a managed tier the operator has already made these choices. You typically get a switch, a destination, and a fixed set of categories. Sometimes allowed decisions are simply not among them — the tier emits connection records and refusals and that is the whole product. Sometimes the granularity is fixed at one level that suits the operator's own cost model rather than your investigation. Neither can be widened by asking nicely in an incident, so the time to find out which you have is before you write a control that depends on it. ## The common mistake The mistake is to switch full recording on in an incident, observe that the cluster copes at that moment, and leave it on. Access traffic is bursty, the trail's cost scales with the busiest hour rather than the average one, and the destination it writes to has its own limits. The safe version is a **time-boxed window**: full recording for a named set of streams, for a stated period, with a decision at the end about which narrower row of the table to settle on.
- Why is writing audit records back onto the same cluster a poor default even ignoring tamper concerns?Because the records are ordinary writes. They add traffic to the cluster you were trying to observe, they are themselves authorized and so generate further decisions unless the audit stream is excluded, and their volume scales with the cluster's busiest hour — exactly when you least want extra load on it.
- What does a sampled audit trail genuinely buy you?Proportions. It shows which principals dominate access to a stream, whether access is spread across many identities or concentrated in one, and when a pattern changes. What it cannot do is answer a question about one named access, because the absence of a sampled record carries no information about whether that access occurred.
- If full recording is affordable on your cluster, is there a reason not to enable it?Only the destination and the buffering. A complete trail is the strongest evidence available, so where the decision is taken once per subscription rather than per request, enabling it is usually right. Check that the destination can absorb the busiest hour, and that whatever is buffered in a node at the moment it dies is an acceptable gap.
saying these in an interview costs you the question
- Treating full audit recording purely as a storage question
- Assuming the per-request cost is the same on every broker design
- Believing a sampled trail can prove one specific access happened
- Leaving full recording on after an incident without re-measuring
- Expecting a managed tier to add a granularity it does not offer