skip to content

Design how a multi-tenant Connect cluster gives each connector its own broker identity for per-tenant ACLs. What pieces fit together?

level: principalimportance: should knowfreq 35%

answer

  1. override.policy=Principal is the linchpin
  2. producer.override./consumer.override.sasl.jaas.config
  3. secrets via ConfigProvider placeholder, not plaintext
  4. per-principal broker ACLs + quotas
  5. worker internal clients keep operator identity (top-level)
  6. REST auth protects the control plane

basics

~10 s

Set connector.client.config.override.policy=Principal on the worker, then give each connector its own credentials via producer.override./consumer.override.sasl.jaas.config (resolved from a ConfigProvider). Define per-principal broker ACLs so each tenant can only touch its own topics.

solid answer

~50 s

Vanilla Connect runs all connectors under the worker's single broker identity, which breaks tenant isolation. The design combines three layers: (1) **Override policy** — set `connector.client.config.override.policy=Principal` so connectors may override only security configs; (2) **Per-connector credentials** — each connector config sets `producer.override.sasl.jaas.config` (source) and/or `consumer.override.sasl.jaas.config` (sink) with that tenant's principal, never literally — pull them via a `ConfigProvider` (`${vault:...}`/`${file:...}`) so secrets stay out of the config topic and REST output; (3) **Broker authorization** — define ACLs (or RBAC) per principal so tenant A's principal can only produce/consume its topics and consumer groups. Complement with REST API auth so tenants can't read each other's configs, and quotas per principal. The worker's own internal clients (config/offset/status) keep a separate operator identity at the top level. This gives least-privilege per tenant with rotation handled by the secret provider.

go deeper

for a junior

Recognize that connectors can use their own credentials and that brokers enforce ACLs per principal.

for a middle

Combine override.policy=Principal with producer.override credentials and per-principal ACLs.

for a senior

Add ConfigProvider externalization, quotas, and REST auth; handle the worker's separate operator identity.

for a principal

Architect full multi-tenant isolation (defense in depth), weigh shared-worker vs cluster-per-tenant, and automate credential/ACL provisioning and rotation.

**The core gap:** A single Connect worker process opens its producer/consumer/admin clients with one set of credentials. Every connector it hosts therefore authenticates to the brokers as the *same* principal. On the broker, ACLs are granted to principals — so if all connectors share one principal, you cannot isolate tenant A's connector from tenant B's topics. Multi-tenant Connect needs each connector to present a *distinct* identity. **Layer 1 — allow per-connector security overrides (worker):** Set `connector.client.config.override.policy=Principal`. With the default `None`, connectors can't override anything. `Principal` permits overriding exactly the security/identity-related client configs (`security.protocol`, `sasl.*`, `ssl.*`) and nothing else — so a connector can change *who it is* but not, say, `bootstrap.servers` (which `All` would dangerously permit). This is the linchpin: it scopes the blast radius of override capability to identity only. **Layer 2 — give each connector its own credentials (connector config):** In each connector's config: ``` producer.override.sasl.jaas.config=org.apache.kafka.common.security.scram.ScramLoginModule required \ username="tenantA" password="${vault:secret/connect/tenantA:password}"; ``` and for sinks the `consumer.override.sasl.jaas.config` equivalent. Crucially the password is a **ConfigProvider** placeholder, not plaintext, so it never lands in the internal config topic or `GET /connectors/{name}/config`. The worker resolves it at start time on each node. **Layer 3 — broker authorization (per-principal ACLs):** On the brokers, grant ACLs (or vendor RBAC) so `User:tenantA` may only Write to tenantA's target topics (source) or Read its source topics + use its consumer group (sink), and nothing else. Now even if a tenant tampers with topic names in their connector config, the broker denies access outside its grant. Add **client quotas** per principal to prevent a noisy tenant starving others. **Layer 4 — protect the control plane (REST):** Per-tenant data isolation is moot if any tenant can read another's connector config or delete it. Secure the REST API (HTTPS + an auth extension or a gateway with RBAC) so tenants only manage their own connectors. ConfigProvider externalization also ensures a tenant reading a config they *can* see doesn't get another's secret. **The worker's own identity:** Keep the worker's internal config/offset/status and group-membership clients on a dedicated *operator* principal set at the **top level** of the worker config (not the producer.override./consumer.override. prefixes, which only affect connector clients). That principal needs ACLs on the internal Connect topics and its group — separate from any tenant. **Edge cases / trade-offs:** - **Resolution per worker:** every worker running a connector must reach the secret store; a missing secret on one node fails that connector there. - **Admin client:** if connectors create topics, also consider `admin.override.*`; otherwise topic creation uses the worker identity. - **Rotation:** prefer a managed provider with TTL so rotating a tenant credential triggers reload without manual edits. - **Operational complexity:** per-connector principals multiply credentials, ACLs, and quotas — automate provisioning (e.g. Terraform/operator) or the toil becomes the failure mode. - **Alternative architecture:** if isolation needs are strong, run a Connect cluster (or worker pool) per tenant instead of multiplexing — simpler blast radius at higher resource cost. The override-policy approach is the lighter-weight multiplexed option. **Why this is the right shape:** Security lives in defense-in-depth — override policy limits what configs change, ConfigProvider keeps secrets out of state, broker ACLs enforce least privilege at the data plane, and REST auth protects the control plane. No single layer is sufficient alone.

  • Why not just use override.policy=All for maximum flexibility?
    All lets connectors override any client config, including bootstrap.servers or disabling TLS hostname verification — a tenant could redirect or weaken the data plane. Principal limits overrides to identity only.
  • If broker ACLs already restrict each principal, why still secure the REST API?
    ACLs protect the data plane, but the REST API is the control plane: without auth a tenant could read another's config (secrets, topics) or delete/modify their connectors. Defense in depth needs both.
  • When would you prefer a Connect cluster per tenant over the override approach?
    When isolation requirements are strict and you want independent blast radius, scaling, upgrades, and failure domains — at the cost of more infrastructure than multiplexing tenants on shared workers.
  • Which principal runs the worker's config/offset/status clients?
    A dedicated operator principal set at the top level of the worker config (not via *.override.*), with ACLs on the internal Connect topics and group — separate from any tenant principal.

saying these in an interview costs you the question

  • Using override.policy=All when Principal suffices (over-broad)
  • Putting tenant passwords as plaintext in connector configs instead of ConfigProvider placeholders
  • Relying on broker ACLs alone while leaving the REST API open
  • Forgetting per-principal quotas, letting one tenant starve others
  • Assuming the worker's internal clients pick up producer.override. settings (they use top-level config)

context