Design an end-to-end audit pipeline that ships Kafka authorization denials and authentication failures to a central SIEM. What sources do you tap and what are the pitfalls?
answer
- two planes: authz log + authn metrics
- Fluent Bit/Filebeat/Vector for logs
- jmx_exporter for metrics
- per-broker decisions, NTP sync
- lock the audit topic; DEBUG volume
basics
~20 sTap two sources: the authorizer audit log (kafka.authorizer.logger, allow/deny lines) and the broker authentication metrics (failed-authentication, expired-connections-killed). Ship logs via a forwarder and metrics via a JMX/Prometheus exporter into the SIEM, normalizing principal, resource, listener, and result.
solid answer
~50 sAn effective Kafka audit pipeline combines log-based and metric-based sources. For ACL decisions you route kafka.authorizer.logger to a dedicated structured appender (DEBUG if you need allows, INFO for deny-only) and forward those lines with an agent (Fluent Bit / Filebeat / Vector) into the SIEM, parsing principal, operation, resource type/name/pattern, host, and Allowed/Denied. For authentication you scrape kafka.server:type=socket-server-metrics — failed-authentication-total, failed/successful-reauthentication, expired-connections-killed-count — through a JMX-to-Prometheus exporter and forward those into the SIEM's metric store or as events. Normalize a common schema (who/what/where/result/timestamp/listener), correlate across brokers (decisions are per-broker — the leader for the partition handles the request), and de-duplicate. Pitfalls: DEBUG allow-logging volume can be enormous; counters reset on restart; super.users bypass logging nuances; clock skew breaks correlation; and you must protect the audit topic/stream itself with ACLs so it can't be tampered with.
go deeper
Know audit data has two sources: the authorizer log and broker auth metrics.
Pick concrete forwarders (log agent + JMX exporter) and the fields to normalize.
Reason about per-broker decisions, counter resets, super-user blind spots, and volume control.
Architect the whole pipeline: canonical schema, tamper resistance, retention/PII, cross-broker correlation, and how it answers real investigations.
## Goal A SIEM (Security Information and Event Management system, e.g. Splunk, Elastic SIEM, Chronicle) needs a faithful, tamper-resistant record of *who did what* on Kafka and *who was rejected*. Kafka surfaces this across two distinct planes, and a good design taps both. ## Source 1 — Authorization (ACL) decisions: the authorizer log - Logger `kafka.authorizer.logger`. **INFO = denials only; DEBUG = allows + denials.** - Route it to its **own appender** with `additivity=false`, ideally emitting **structured/JSON** lines (custom Log4j2 layout) so the SIEM parses fields rather than regexing free text. - Fields to capture: principal (`User:...`), operation, resource (type, name, LITERAL/PREFIXED pattern), host, and result (Allowed/Denied), plus broker id and timestamp. - Forward with a log agent: **Fluent Bit, Filebeat, or Vector** tailing the audit file → SIEM ingest. ## Source 2 — Authentication outcomes: broker metrics - MBean `kafka.server:type=socket-server-metrics` per listener: `failed-authentication-total/-rate`, `successful-authentication-total`, `failed-reauthentication-total`, `expired-connections-killed-count`. - Expose via the **Prometheus jmx_exporter** Java agent, then forward metrics (or derived events/alerts) into the SIEM. - These tell you about bad creds, expired certs/tokens, brute-force patterns — none of which appear in the authorizer log. ## Normalization & correlation - Define one **canonical event schema**: timestamp, broker, listener, principal, action/operation, resource, result, source-host. - **Per-broker reality:** an authorization decision is made on whichever broker handles the request (e.g. the partition leader). The same principal hits different brokers, so you must aggregate cluster-wide and key correlation on principal + time, not on a single broker. - **Clock discipline:** brokers must be NTP-synced or cross-broker correlation and ordering in the SIEM break. ## Pitfalls (the senior judgment) 1. **Volume:** DEBUG allow-logging on a busy cluster can dwarf application logs. Options: deny-only at INFO in steady state; sample allows; or only DEBUG specific principals/resources during investigations. 2. **Counter resets:** `-total` metrics reset on restart — use rate/increase, and don't treat a drop as 'fixed'. 3. **Super users bypass:** `super.users` are allowed without ACL evaluation; ensure their actions are still captured (they are logged as allowed at DEBUG) so privileged activity isn't a blind spot. 4. **allow.everyone.if.no.acl.found:** if true, missing-ACL access is allowed and only visible at DEBUG — a stealthy over-permission risk. 5. **Tamper resistance:** if you pipe audit events *through Kafka itself*, lock the audit topic with strict ACLs and ideally a separate cluster, so a compromised principal can't delete its own trail. 6. **Gaps:** the authorizer log doesn't record authentication; the auth metrics don't record which resource was targeted. You need *both* for a complete story. 7. **Cardinality/PII:** principal and host fields may be sensitive; apply retention and access controls in the SIEM. ## Result The SIEM can then answer: 'show every Denied Write to Topic payments in the last hour', 'alert on failed-authentication spikes on the EXTERNAL listener', and 'correlate a burst of denials with a credential that just expired' — which no single Kafka source could answer alone.
- Why isn't the authorizer log alone sufficient for a complete audit trail?It only records authorization (ACL allow/deny) and never authentication. Failed logins, SSL handshake failures, and expired-credential disconnects appear only in the socket-server-metrics, so you must combine both sources to see both 'rejected at the door' and 'rejected by policy' events.
- If you transport audit events through a Kafka topic, how do you keep them tamper-resistant?Restrict the audit topic with tight ACLs (only the audit producer can write, only the forwarder can read, no delete/admin for normal principals), ideally place it on a separate cluster, and forward promptly to immutable SIEM storage so a compromised principal can't erase its own trail.
- Why must brokers be NTP-synchronized for this pipeline?Authorization decisions and auth metrics are produced independently on each broker. Without synchronized clocks, the SIEM can't correctly order or correlate events across brokers, breaking timeline reconstruction and spike detection.
saying these in an interview costs you the question
- Relying only on the authorizer log and ignoring authentication metrics (or vice versa).
- Leaving DEBUG allow-logging on globally without considering volume.
- Forgetting that super.users and allow.everyone.if.no.acl.found create audit blind spots.
- Shipping audit through an unprotected Kafka topic that the monitored principals can delete.