How would you structure Kafka alerting around SLOs to avoid alert fatigue — symptom-based vs cause-based alerts, escalation tiers, and what should and shouldn't page?
answer
- Design top-down from SLOs, not bottom-up from metrics
- Page on symptoms, diagnose with causes
- SLOs: availability, durability, freshness
- Tiers: page / ticket / dashboard
- Burn-rate: fast burn pages, slow burn tickets; inhibit + group by root cause
basics
~20 sPage on a small set of symptom alerts tied to user-facing SLOs (availability, durability, freshness) — offline partitions, write failures, SLO-breaching lag. Route cause-based and predictive signals (disk filling, URP, rising controller queue) to tickets/dashboards, not pages, unless they breach an SLO.
solid answer
~50 sAnchor alerting to a few SLOs that mirror what users feel: availability (can clients produce/consume?), durability (are we keeping the promised replicas/acks?), and freshness (is consumer lag within target?). Then split alerts: symptom-based alerts page because they mean an SLO is breaching now — offline partitions, UnderMinIsr (acks=all failing), SLO-breaching consumer lag, no active controller. Cause-based/diagnostic alerts (URP without min-ISR breach, disk at 70%, rising controller event queue, elevated request latency) go to tickets and dashboards; they're how you fix incidents and prevent them, but they shouldn't wake someone unless they're predictive of imminent SLO breach (e.g. disk projected full within the on-call window). Use multi-tier severity: page for SLO-breaching, ticket for degraded, dashboard for informational. Add error-budget-style burn-rate alerting for lag/availability so slow burns ticket and fast burns page. Always require sustained windows, inhibit during maintenance, and group alerts by root cause so one broker failure is one incident, not 200 pages.
go deeper
Understand that not every metric should page and that user-facing problems get priority.
Separate page-worthy symptoms from ticket-worthy causes and apply sustained windows.
Design SLO-anchored tiers, burn-rate alerts, and inhibition/grouping so one failure is one incident.
Own the SLO/error-budget framework and escalation policy org-wide, negotiate targets with consumers, and teach symptom-vs-cause discipline to reduce fatigue while keeping diagnostic coverage.
## The problem: alert fatigue Kafka exposes hundreds of metrics. If you page on each one nonzero, one broker failure fires dozens of simultaneous pages, on-call burns out, and real signals get ignored. The fix is to design alerting top-down from **SLOs** (what users are promised) rather than bottom-up from available metrics. ## SLOs for Kafka Map to what consumers of the platform actually experience: - **Availability**: producers and consumers can read/write — proxied by offline partitions = 0 and acceptable request error rates. - **Durability**: the cluster honors its replication/acks contract — proxied by ISR >= min.insync.replicas (no UnderMinIsr) and zero unclean leader elections. - **Freshness (for streaming consumers)**: records processed within a target latency — proxied by consumer lag-in-seconds under an SLO. Each SLO gets a target (e.g. 99.95% of minutes with zero offline partitions) and an **error budget** (the allowed shortfall). ## Symptom-based vs cause-based alerts - **Symptom-based** (page-worthy): describe a user-visible failure happening *now*. Offline partitions, UnderMinIsrPartitionCount > 0, consumer lag past the freshness SLO, ActiveControllerCount sum != 1, sustained produce/fetch error rates. These page because the SLO is breaching. - **Cause-based / diagnostic** (ticket/dashboard): explain *why* something might break — URP without a min-ISR breach, disk at 70-80%, rising controller event-queue size, GC pause time, request-handler thread-pool idle ratio dropping. These are essential for diagnosis and prevention but are not themselves user-visible failures. Rule of thumb: **page on symptoms, diagnose with causes.** A cause-based alert is promoted to a page only when it's *predictive of an imminent SLO breach* — e.g. disk fill rate projects 100% within hours (a partition going offline soon), or lag burn-rate is steep enough to blow the freshness budget within the shift. ## Escalation tiers 1. **Page (critical)** — SLO breaching or about to: offline partitions, write failures, controller absent, fast lag burn. Wakes someone. 2. **Ticket (warning)** — degraded, fix during business hours: sustained URP, disk 70-80%, slow controller queue, elevated latency, slow lag burn. 3. **Dashboard / info** — trends and capacity: throughput, partition counts, ISR churn history. No notification. ## Burn-rate / error-budget alerting Borrowed from SRE: alert on how fast you're consuming the error budget. A **fast burn** (budget gone in hours) pages; a **slow burn** (budget gone in days) tickets. This naturally separates urgent from chronic without per-metric threshold guessing, and is ideal for lag and availability. ## Noise-reduction mechanics - **Sustained windows** ('for 10m') on everything page-worthy except hard outages, so transients don't fire. - **Inhibition / grouping**: when a broker is down, suppress the 200 downstream URP/leadership alerts and raise one 'broker X down' incident. Alertmanager-style inhibition rules or a higher-severity alert that mutes its children. - **Maintenance silences** around rolling restarts and reassignments. - **Dedup by root cause**: alert on counts and the common dimension (broker, AZ), not per-partition. ## Runbook linkage and escalation policy Every page-worthy alert links to a runbook and an escalation path (on-call → data-platform owner → vendor). This is what turns a metric into an actionable incident response and keeps MTTR low. ## Edge cases / tradeoffs - Too few alerts and you miss slow-burn degradations (e.g. lag creeping toward the SLO); burn-rate alerting covers this. - Symptom-only paging can leave you blind to *why*; that's exactly why cause-based alerts still exist — just routed to tickets/dashboards. - SLO targets must be set with the consumers of the platform, not invented by ops, or the error budget is meaningless.
- Give an example of a cause-based alert that should normally ticket but sometimes page.Disk usage on log.dirs. At 70% it tickets (capacity work). But if the fill-rate projects 100% within the on-call window, it's predictive of partitions going offline — promote it to a page, because an imminent SLO breach justifies waking someone even though the cause itself isn't yet user-visible.
- How does burn-rate alerting help with consumer lag specifically?Instead of a fixed lag threshold, you frame freshness as an SLO with an error budget and alert on how fast lag is consuming it. A steep burn (budget gone within the shift) pages as an active incident; a shallow burn (budget gone over days) tickets for tuning — separating urgent from chronic without guessing absolute numbers.
saying these in an interview costs you the question
- Paging on every nonzero metric, causing one broker failure to fire dozens of pages
- Eliminating cause-based alerts entirely (you lose diagnostic and predictive signal)
- Setting SLO targets in ops without the platform's consumers
- No alert grouping/inhibition, so symptoms and downstream effects all page separately