skip to content

Broker Security Controls

Who may connect to a broker, what each principal may do to which stream, and what is encrypted in flight. Asked because a cluster is usually opened to a whole organisation once and tightened never.

part ofBroker & streaming operationsoverview, primer and where to startread it →
on this pageshow

questions

25

What does a single broker grant have to name before the cluster can decide whether a request is allowed?

level: juniorimportance: must knowfreq 70%

answer

  1. a lookup, not a login
  2. one rule, several columns
  3. identity, verb, named target
  4. plus permit or refuse
  5. and what an unmatched request gets

basics

~20 s

A grant names four things: the principal making the request, the operation attempted, the named resource it is attempted on, and whether that combination is permitted or refused. Leave any one loose and the rule decides more than intended.

solid answer

~50 s

Authorization on a broker is a lookup, not a login. By the time a request arrives the cluster already holds a `principal` for the connection, whatever mechanism established it; the remaining question is whether anything in the grant table binds that principal to this operation on this named resource, with an effect of permit or refuse. Those four parts are the whole unit, and each one is a place where a rule quietly widens: a shared principal used by six services cannot be narrowed to one of them, a coarse operation stands in for writing and deleting alike, and a resource written as a name fragment covers names nobody has created yet. What happens when nothing matches varies by platform - most refuse once enforcement is switched on, while several ship permissive until an operator enables it, so "we have grants" and "the grants are enforced" are two different claims.

go deeper

for a junior

Recall the four parts in order - which identity, doing what, to which named resource, permitted or refused - and remember that connecting successfully is a separate step from being allowed to act.

for a middle

Explain how each part widens in practice: a shared identity, a coarse operation, a name fragment as the resource. Say that the unmatched-request default varies between platforms rather than asserting one.

for a senior

Demonstrate that you check enforcement, not just the table: whether authorization is switched on, what an unmatched request actually gets, and whether a refusal is distinguishable from a missing stream.

for a principal

Frame the four parts as the estate's vocabulary. If teams cannot describe a rule in those terms, grants will be issued per team rather than per operation, and no cluster in the estate will be comparable to another.

## The question a broker is actually answering When a client asks a broker to append a record, read from a stream, or create one, the cluster is not asking *who are you* - that was settled when the connection was established and left the broker holding a **principal**, an identity string it can compare against rules. Every later request reduces to one lookup: **is there anything in the grant table that lets this principal perform this operation on this named resource?** A **grant** is the unit that makes that lookup possible, and the reason interviewers ask about its shape is that each of its parts is a place where real systems leak. ## The four parts - **The principal** - the identity the broker holds after the connection is established. It may have come from a password-style exchange, a certificate, an issuer-signed short-lived token, or from the platform the client runs on vouching for it; the grant does not care which. It cares that the identity is *specific enough to narrow*. A single identity shared by six deployments is a ceiling on how tight any rule can ever be. - **The operation** - the verb being attempted. Brokers distinguish far more than reading and writing, and a grant that names a coarse operation is a grant to everything that coarse operation covers. - **The resource** - the named thing the operation acts on: a stream, a namespace, sometimes the cluster itself. A resource can be written as an exact name, as a **prefix grant** over a name fragment, or as a match-all. The last two are evaluated against whatever names exist at the moment of the request, not the names that existed when the rule was written. - **The effect** - whether this combination is permitted or refused. Most grant tables are written entirely in permits, with refusal being the absence of a match; some platforms also let you write an explicit refusal. | Part | What it answers | What goes wrong when it is left loose | |---|---|---| | Principal | who is asking | one shared identity behind several workloads, so no rule can be narrowed to one of them | | Operation | what they are attempting | one broad right standing in for appending, creating, altering and deleting | | Resource | on what name | a fragment or match-all that reaches names created long afterwards | | Effect | permit or refuse | rules written as documentation that nothing actually evaluates | ## What happens when nothing matches This is where platforms genuinely differ, and a candidate who states one behaviour as universal is describing the product they happen to know. Designs vary along two axes worth naming: 1. **The default.** Most brokers refuse anything no grant covers once authorization is enabled - but on several, enforcement is itself something an operator turns on, and until then every request is permitted regardless of what the table says. A cluster can therefore have a carefully written grant table and be wide open. 2. **What the client is told.** Some platforms report a refusal distinctly from a missing stream; others answer identically on purpose, so that an unauthorised caller cannot map what exists. A client-side error message is evidence, not proof. ## Where the rule lives, and who evaluates it The four parts survive either arrangement. The broker may hold its own grant table and evaluate it itself, or it may delegate the decision to the surrounding platform, which evaluates its own rules and answers permit or refuse. A rented cluster may only offer you the second. Either way the same four questions are being asked; what changes is who can change a rule, how finely the operations can be expressed, and what happens when the deciding system is unreachable. ## Why this is a first-screen question Because the failures it predicts are the ones interviewers have lived through. A team that cannot say what a grant names will write one grant per team rather than per operation, will hand out a fragment resource because it saves a ticket, and will read a short grant table and conclude the cluster is tight. Knowing the shape is what turns "we have security on the cluster" into four checkable questions: *which identity, doing what, to which names, permitted or refused.*

  • If the grant table is empty, is the cluster open or closed?
    It depends on the platform, and that is the point. Most brokers refuse anything no grant covers once authorization is enforced, so an empty table means nobody but the unrestricted setup identity can do anything. But on several platforms enforcement is a switch an operator must throw, and until it is thrown the table is inert and every request succeeds. Confirm which state the cluster is in rather than inferring it from the rules.
  • Why is a grant written over a client's network address a weaker rule than one written over a principal?
    An address says where a connection came from, not who is on it. Addresses are reassigned, shared behind a gateway, and reused by the next workload scheduled onto the same host, so the rule silently transfers to whoever inherits the address. A principal is derived from a credential the caller had to present, so the rule follows the identity rather than the location. Address rules are useful as a coarse extra filter, never as the identity.

saying these in an interview costs you the question

  • Says a successful connection means the request will be allowed
  • Treats reading and writing as the only operations a grant can name
  • Assumes every broker permits anything no grant mentions
  • Writes grants over a network address instead of a principal
  • Believes a rule that exists in the table is necessarily being enforced
open as a page

When a client opens a connection to a broker, what families of proof can it present, and what does each one establish?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A client proves itself with a shared secret through a salted challenge-response password exchange, with a client certificate, with a directory-issued or issuer-signed short-lived token, or with delegated platform identity where it holds no secret at all.

open as a page

The volume beneath a broker cluster is encrypted at the storage layer — which reader does that stop, and which does it not?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Volume-level encryption beneath the broker stops bytes that leave the cluster: a removed drive, a copied volume image, replaced hardware. It stops nobody who connects, because the storage layer decrypts beneath the broker process, which then serves plaintext to any principal holding a grant.

open as a page

A broker cluster encrypts client connections but not the node-to-node hop — what is exposed there?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Whole record payloads, their keys and the cluster's coordination traffic. Where copies are kept by nodes forwarding records to each other, that hop carries every record again for each extra copy, continuously, for as long as the cluster runs.

open as a page

Beyond reading and writing records, which operations does a broker authorize separately, and why does that separation matter?

level: middleimportance: must knowfreq 60%

basics

~20 s

Brokers separate far more than reading and writing: creating a stream, deleting it, changing its settings, listing or describing names, recording a read position, and cluster-wide administration are distinct rights. Teams that grant only two end up granting all of them.

open as a page

A cluster demands a credential on its public client address but accepts any connection on another of its addresses — how?

level: middleimportance: must knowfreq 65%

basics

~20 s

Authentication is configured per address, not per cluster. A broker answers on several separately-configured endpoints — a public client edge, an internal or administrative one, and the node-to-node hop — and the internal ones are the ones nobody went back to tighten.

open as a page

During a broker credential swap with traffic flowing, what must be true of every node before any client moves to the new credential?

level: middleimportance: must knowfreq 66%

basics

~20 s

Every node must already accept both the outgoing and the incoming credential. Clients reconnect to whichever member answers, so a trust set widened on only some nodes produces intermittent refusals rather than a clean, visible break.

open as a page

Your broker's audit records land on a stream in the same cluster, writable by the principals they describe. What does that cost you?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Two properties. A principal with write access to the trail can edit the account of itself, so the records stop being evidence; and the trail dies with the cluster it describes. Ship it off-cluster, append-only, under different custody.

open as a page

Before withdrawing a broker's old credential, what evidence shows nothing still presents it, and what will that evidence miss?

level: seniorimportance: must knowfreq 52%

basics

~20 s

The cluster's own connection and authorization-decision records, read over an observation window longer than the slowest client's reconnect interval, and only if the two credentials are distinguishable in them. They miss every holder that has not connected during that window.

open as a page

A broker's audit trail is switched on. What kinds of record does it hold, and which question can it usually not answer?

level: juniorimportance: should knowfreq 44%

basics

~20 s

A broker's audit trail holds connection records — who connected, from where, on which endpoint — and authorization-decision records. Most deployments keep only the refusals, so the trail rarely shows who successfully read a stream.

open as a page

Why does a broker credential reaching its expiry date take a cluster down all at once rather than degrading gradually?

level: juniorimportance: should knowfreq 54%

basics

~20 s

Because expiry is a date shared by everything issued in the same batch: every holder loses the right to connect at the same moment. Where the node-to-node hop uses material from that batch, replication stops alongside the clients.

open as a page

Recording every allowed authorization decision on a busy broker is usually rejected. What drives that cost, and what narrower configurations keep some value?

level: middleimportance: should knowfreq 38%

basics

~20 s

The decision sits on the request-serving path, so a complete trail costs work per authorized operation and produces volume comparable to the traffic itself. Narrower options: refusals only, a named set of sensitive streams or principals, one record per session, or sampling.

open as a page

Why does a prefix or wildcard grant written in a hurry end up covering streams nobody had created when it was written?

level: middleimportance: should knowfreq 62%

basics

~20 s

Because a grant whose resource is a name fragment is matched at request time against whatever names exist then, not against the names that existed when it was written. Every stream created later under that fragment is covered the moment it appears.

open as a page

Identity is settled when a broker connection is opened, so what does that mean for a client holding a short-lived signed token?

level: middleimportance: should knowfreq 52%

basics

~20 s

Authentication is an event, not a state: a long-lived connection keeps the principal it got at connect. Some designs close it at expiry, some re-authenticate in place, some never re-check — so expiry usually bites at the next reconnect.

open as a page

When the writing application encrypts each payload before it reaches the broker, whose access changes compared with encrypting the volume underneath?

level: middleimportance: should knowfreq 54%

basics

~20 s

The cluster's own access changes. Volume-level encryption leaves the broker and every grant holder reading plaintext; payload encryption by the writer leaves the cluster holding ciphertext it has no encryption key for, so the threat it answers is the operator and the grant table themselves.

open as a page

A team calls records "encrypted end to end" because every broker connection is encrypted — where does that protection actually stop?

level: middleimportance: should knowfreq 48%

basics

~20 s

At every node the record passes through. An encrypted connection protects one leg; the node decrypts to place, append and retain the record, and re-encrypts it for each reader later — so the protection has gaps wherever the cluster holds the bytes.

open as a page

Why does a cluster's grant table not show what the bootstrap administrator it was set up with can do?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Because the bootstrap administrator is checked before the grant table, not against it. The cluster skips evaluation entirely for that identity, so reading the rules tells you nothing about the one principal that can do everything.

open as a page

A broker accepts clients that present no secret because the platform they run on vouches for them — what is actually being trusted?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The broker trusts the platform's signer, not the client. It checks a short-lived signed statement about the workload — who signed it, that it is current, that it names this cluster — and derives the principal from it.

open as a page

A broker's old credential was withdrawn on Monday with no errors, yet a service redeployed on Wednesday cannot connect — why?

level: seniorimportance: should knowfreq 46%

basics

~20 s

That service's connection predated the withdrawal and was never re-authenticated, so it kept working on a credential the cluster no longer accepts. The redeploy forced a fresh connection, which is the first moment the withdrawal was actually tested for it.

open as a page

A cluster now holds only ciphertext because every writer encrypts its own payloads — what has the cluster permanently given up?

level: seniorimportance: should knowfreq 44%

basics

~20 s

It has given up everything that depends on reading inside a record: selection by content, cluster-side transforms, content-derived routing, and any inspection during an incident. What survives is whatever sits outside the ciphertext — the partitioning key, sizes, timestamps and counts — and that is now the only handle anyone has.

open as a page

After a cluster's connections were encrypted, readers replaying old history slowed far more than caught-up readers — why?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Encryption forces the node to handle bytes it previously never touched, closing the copy-free read path from stored file to network. The added work scales with bytes shipped, and a history replay ships the most bytes fastest; a caught-up reader trickles.

open as a page

You must withdraw broker grants nobody uses, but the audit trail records only refusals. How do you establish use?

level: principalimportance: should knowfreq 33%

basics

~20 s

A refusal-only trail is negative evidence and cannot show use. Prune by connection records first, measure positively for a time-boxed window on the streams under review, then withdraw in stages — and make grants expire so the burden shifts to justifying renewal.

open as a page

Should grants over streams be evaluated by the broker itself or by the platform the cluster runs on, and what does each cost?

level: principalimportance: should knowfreq 45%

basics

~20 s

The broker's own evaluation expresses stream operations precisely and survives the platform being unreachable, but is a second permission system to run. Platform evaluation unifies the estate and loses granularity. Most large estates end up with both.

open as a page

Across an estate of hundreds of client applications on several platforms, how do you decide which authentication mechanisms a broker accepts?

level: principalimportance: should knowfreq 40%

basics

~20 s

Decide on first delivery, failure shape, identity granularity and reach rather than abstract strength: how a new client gets its credential with nobody copying it, whether failure is a quiet leak or a scheduled outage, and which clients can implement it.

open as a page

Payloads on a long-retained stream were encrypted by their writers, so a year on, what does custody of those encryption keys decide?

level: principalimportance: should knowfreq 33%

basics

~20 s

Custody decides who can read the stream's history, including records written long before the current custodians arrived. It is an access decision and a durability decision at once: withdraw or lose the encryption key and the bytes survive as bytes nobody can open, whatever the grant table says.

open as a page