skip to content

MM2 Heartbeats and Replication Monitoring

Using heartbeat topics and replication-lag metrics to prove a mirror is alive and to measure real RPO. Interviewers ask how you would notice that replication had silently stopped.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the MirrorHeartbeatConnector in MirrorMaker 2, and what is the heartbeats topic used for?

level: juniorimportance: must knowfreq 55%

answer

  1. 3 connectors: Source / Checkpoint / Heartbeat
  2. writes 'heartbeats' topic on source
  3. record carries source alias + target alias + timestamp
  4. interval = emit.heartbeats.interval.seconds (1s)
  5. liveness + end-to-end latency probe

basics

~10 s

MirrorHeartbeatConnector periodically writes timestamped records to a 'heartbeats' topic on the source cluster. MM2 replicates them to the target, proving the replication path is alive and letting you measure how fast it flows.

solid answer

~40 s

MM2 has three connectors: MirrorSourceConnector (copies data topics), MirrorCheckpointConnector (translates consumer offsets), and MirrorHeartbeatConnector. The heartbeat connector runs against the source cluster and, on a fixed interval (emit.heartbeats.interval.seconds, default 1s), produces small records into a topic named 'heartbeats' containing the source cluster alias, target alias, and a timestamp. MirrorSourceConnector then replicates that heartbeats topic to the target (as <source>.heartbeats). Because the records carry the emit timestamp, a consumer on the target can compare it to wall-clock time and compute end-to-end replication latency, and the steady arrival of new heartbeats proves the whole flow (source produce -> mirror -> target) is alive. It's the canonical liveness and lag probe for a replication flow, independent of whether real data topics are currently active.

go deeper

for a junior

Know there are 3 MM2 connectors and that the heartbeat one writes a 'heartbeats' topic to prove the link is alive.

for a middle

Explain the record contents (aliases + timestamp) and how it enables both liveness and latency measurement.

for a senior

Tie heartbeats to RPO observation and explain replication of the topic to <source>.heartbeats on the target.

for a principal

Reason about heartbeats as the topology-wide health/SLA signal and how to alert on heartbeat staleness across many flows.

## Background MirrorMaker 2 (MM2) is Kafka's built-in cross-cluster replication tool, built on Kafka Connect. It runs three connector types per replication 'flow' (a direction such as A->B): - **MirrorSourceConnector** — copies records of matched topics from source to target, renaming them with the source alias prefix (e.g. topic `orders` from cluster `us-east` becomes `us-east.orders` on the target by default). - **MirrorCheckpointConnector** — emits offset checkpoints so a consumer can fail over and resume near where it left off. - **MirrorHeartbeatConnector** — the subject here. ## What MirrorHeartbeatConnector does It is a tiny *source connector* that does not read any real data. On a fixed schedule it **produces synthetic records** into a topic literally named `heartbeats` on the **source** cluster. Each record's value/key encodes: the **source cluster alias**, the **target cluster alias**, and a **timestamp** (the moment the heartbeat was emitted). The interval is controlled by `emit.heartbeats.interval.seconds` (default **1 second**). Whether heartbeats are emitted at all is gated by `emit.heartbeats.enabled` (default **true**). ## Why a separate heartbeats topic Once `heartbeats` exists on the source, MirrorSourceConnector replicates it like any other topic, producing `<source-alias>.heartbeats` on the target. So the heartbeat record travels the **exact same path** real data takes. That gives you two things for free: 1. **Liveness** — if new heartbeat records keep arriving on the target, the entire pipeline (source broker is reachable, MM2 Connect workers are running, target broker is accepting writes) is healthy. A *stall* in heartbeat arrivals is an early, data-independent alarm even when no business topics are being written. 2. **End-to-end latency / RPO observation** — because each record carries its emit timestamp, a monitor reading the target can compute `now - heartbeat.timestamp` to estimate how far behind the target is. That delta is a practical proxy for **Recovery Point Objective (RPO)**: how much data you'd lose if the source vanished right now. ## Edge cases - Heartbeats are independent of data traffic, so they detect a broken flow even on idle topics. - In bidirectional setups both directions emit their own heartbeats (`A.heartbeats` on B, `B.heartbeats` on A). - Heartbeat volume is tiny but non-zero; on very large fan-out topologies the 1s default is sometimes relaxed to a few seconds. - Disabling heartbeats removes your simplest liveness/latency probe — generally discouraged.

  • What does a heartbeat record actually contain?
    The source cluster alias, the target cluster alias, and the emit timestamp — enough to identify the flow and measure latency on arrival at the target.
  • If real data topics are idle, can heartbeats still tell you the flow is healthy?
    Yes. Heartbeats are synthetic and emitted on a schedule regardless of data traffic, so they keep flowing and prove liveness even when no business topic is being produced to.

saying these in an interview costs you the question

  • Saying heartbeats are produced on the target cluster (they originate on the source, then are replicated)
  • Claiming the heartbeat connector reads/echoes real data — it produces synthetic records only
  • Confusing MM2 heartbeats with Kafka consumer-group heartbeats to the group coordinator — unrelated mechanisms

context

open as a page

Which MM2 JMX metrics would you use to monitor replication lag and latency, and what does each one mean?

level: seniorimportance: must knowfreq 45%

basics

~10 s

MM2's MirrorSourceConnector exposes per-topic-partition JMX metrics: replication-latency-ms (time from source append to target append), record-age-ms (age of a record when consumed from source), and byte/record rates. You scrape these to track lag.

open as a page

Explain emit.heartbeats.interval.seconds and the related heartbeat configs. How would you tune them and what are the trade-offs?

level: middleimportance: should knowfreq 40%

basics

~10 s

emit.heartbeats.interval.seconds sets how often MirrorHeartbeatConnector writes a heartbeat record (default 1s). emit.heartbeats.enabled (default true) turns the feature on/off. Shorter interval = finer latency resolution but more overhead.

open as a page

How would you use MM2 heartbeats to observe end-to-end replication latency and estimate RPO for a DR setup?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Consume the replicated <source>.heartbeats topic on the target, read each record's emit timestamp, and compute now - timestamp. That delta is your end-to-end replication latency, which approximates worst-case RPO — how much data you'd lose on failover.

open as a page

Heartbeats have stopped arriving on the target for a flow, but the MM2 Connect cluster looks 'up'. How do you diagnose this, and what does heartbeat staleness tell you versus rising replication-latency-ms?

level: principalimportance: should knowfreq 25%

basics

~20 s

Stale heartbeats mean the flow is broken or stalled, not merely slow. Check the heartbeat connector/task state, the source heartbeats topic, and the MirrorSourceConnector replicating it. Rising replication-latency-ms instead means the flow works but is lagging.

open as a page