Explain emit.heartbeats.interval.seconds and the related heartbeat configs. How would you tune them and what are the trade-offs?
answer
- emit.heartbeats.interval.seconds default = 1s
- emit.heartbeats.enabled default = true
- interval = sampling rate of liveness/latency
- floor is 1s (no sub-second)
- raise it only for large fan-out topologies
basics
~10 semit.heartbeats.interval.seconds sets how often MirrorHeartbeatConnector writes a heartbeat record (default 1s). emit.heartbeats.enabled (default true) turns the feature on/off. Shorter interval = finer latency resolution but more overhead.
solid answer
~40 sThe heartbeat connector is governed mainly by two configs. `emit.heartbeats.enabled` (default true) controls whether heartbeats are produced at all. `emit.heartbeats.interval.seconds` (default 1) controls the cadence — every N seconds a new record lands in the source `heartbeats` topic. The interval sets the *resolution* of your latency and liveness signals: with a 1s interval you can detect a stall and estimate latency to ~1s granularity. Tuning down (sub-second isn't supported; minimum is 1s) gives finer signal at the cost of more produce traffic and Connect task work across every flow. Tuning up (e.g. 5–10s) reduces overhead in large topologies but blunts your detection latency and RPO estimate. A related setting, `replication.policy.separator` and the heartbeats topic naming, affects where heartbeats land on the target (`<source>.heartbeats`). Most teams keep the 1s default unless they run many flows.
go deeper
Know the config name, the 1s default, and that it controls how often heartbeats are written.
Articulate the interval as a sampling-rate trade-off and distinguish it from the checkpoint interval.
Discuss tuning for large topologies and aligning staleness-alert thresholds to the chosen interval.
Frame heartbeat cadence within an SLA/RPO budget across the whole replication topology and its overhead model.
## The configs - **`emit.heartbeats.enabled`** — boolean, default **true**. When true, the MM2 herder runs a MirrorHeartbeatConnector for the flow and it produces heartbeat records. Setting it false stops all heartbeat emission, removing your cheapest liveness/latency probe. - **`emit.heartbeats.interval.seconds`** — integer seconds, default **1**. Period between heartbeat records written to the source `heartbeats` topic. - (Related, often confused) **`emit.checkpoints.interval.seconds`** — default 60 — governs the *checkpoint* connector, not heartbeats. Don't mix them up. These are set per-flow in the MM2 properties, e.g. `A->B.emit.heartbeats.interval.seconds = 5`, or as defaults across flows. ## What the interval buys you The heartbeat interval is the **sampling rate** of your replication-health signal: - **Liveness detection latency** — if you alert on 'no new heartbeat for X seconds', X must be a multiple of the interval. A 1s interval lets you alert quickly; a 30s interval means you can't notice a stall faster than ~30s. - **Latency/RPO resolution** — each heartbeat is a sample of `now - emitTimestamp`. More frequent samples = smoother, more current latency curve and a tighter RPO estimate. ## Trade-offs of tuning - **Shorter (toward 1s, the floor):** finer signal, faster detection. Cost: more produce requests, more records to replicate, more Connect task scheduling — multiplied by the number of flows. On a handful of flows this is negligible. - **Longer (5–30s):** materially less overhead when you have dozens/hundreds of flows or are bandwidth-sensitive. Cost: coarser detection and a laggier RPO view. You might also miss brief stalls entirely. ## Edge cases & gotchas - The practical minimum is **1 second**; you cannot get sub-second heartbeats this way. - Disabling heartbeats (`enabled=false`) is occasionally done to cut noise on huge topologies, but then liveness must come from another signal (e.g. replication-latency JMX on data topics, broker-side metrics). - The heartbeats topic itself is tiny; the cost is dominated by per-record connector/produce overhead across many flows, not storage. - Changing the interval doesn't retroactively change historical samples — it only affects the cadence going forward. ## Rule of thumb Keep the 1s default for a few flows. Raise it (e.g. 5s) only when heartbeat overhead is measurably significant in a large fan-out topology, and re-tune your staleness alert thresholds to match.
- What's the smallest heartbeat interval you can configure, and what does that bound your detection at?1 second is the practical floor, so you can't sample replication health (or estimate latency/RPO) at finer than ~1s granularity through heartbeats.
- Someone set emit.heartbeats.interval.seconds=60 — what monitoring consequence should you call out?Detection of a stalled flow is now coarse (~60s), and the latency/RPO estimate refreshes only once a minute, so brief outages may go unnoticed.
saying these in an interview costs you the question
- Claiming sub-second heartbeat intervals are configurable
- Confusing emit.heartbeats.interval.seconds with emit.checkpoints.interval.seconds (default 60)
- Saying a shorter interval reduces overhead — it increases it
- Asserting disabling heartbeats has no monitoring downside