skip to content

Heartbeats have stopped arriving on the target for a flow, but the MM2 Connect cluster looks 'up'. How do you diagnose this, and what does heartbeat staleness tell you versus rising replication-latency-ms?

level: principalimportance: should knowfreq 25%

answer

  1. staleness = binary stall; latency = continuous lag
  2. 'Connect up' != task up — check per-task REST status
  3. source heartbeats fresh? isolates emit vs replication leg
  4. dead task => metrics absent, not zero
  5. rule out ACL/auth, rebalance storm, clock skew

basics

~20 s

Stale heartbeats mean the flow is broken or stalled, not merely slow. Check the heartbeat connector/task state, the source heartbeats topic, and the MirrorSourceConnector replicating it. Rising replication-latency-ms instead means the flow works but is lagging.

solid answer

~50 s

Heartbeat staleness (no new <source>.heartbeats records on the target) is a *binary* failure: something in the path is stopped. Rising replication-latency-ms is *continuous* — the path works but is behind. Diagnose staleness by checking, in order: (1) is MirrorHeartbeatConnector and its task RUNNING on the source side (Connect REST /connectors/.../status)? A FAILED/paused task stops emission; (2) is the source `heartbeats` topic actually receiving new records (consume it on the source)? (3) is MirrorSourceConnector — which replicates heartbeats to the target — RUNNING, or has it failed/rebalanced? (4) target broker reachability/ACLs/topic auto-create for `<source>.heartbeats`. 'Connect is up' only means workers are alive, not that the specific tasks for this flow are. Also rule out: dead task emitting no metrics, an offset/rebalance storm, or consumer-side auth failures. Heartbeat staleness is your DR-not-ready alarm; replication-latency-ms is your performance/SLA gauge.

go deeper

for a junior

Know that no new heartbeats means the replication is broken, not just slow.

for a middle

Check connector/task status and the source heartbeats topic to localize the problem.

for a senior

Distinguish binary staleness from continuous latency, isolate emit vs replication leg, and consider ACL/rebalance causes.

for a principal

Build the full diagnostic decision tree and alerting policy: per-task REST checks, metric-absence handling, clock-skew/auth edge cases, and DR-readiness escalation.

## Two fundamentally different signals - **Heartbeat staleness** = *no new heartbeat records arriving*. This is **binary**: the replication path for that flow is **stopped or broken**. It is the strongest DR-readiness alarm because if heartbeats can't get through, real data can't either. - **Rising `replication-latency-ms`** = records *are* flowing but the source-append-to-target-append delay is growing. This is **continuous degradation** — a performance/SLA problem, not an outage. Confusing the two leads to wrong responses: you page for an outage when it's just lag, or you treat a true stall as 'a bit slow'. ## Why 'Connect is up' is misleading MM2 runs as Connect connectors/tasks. The Connect **worker JVMs** can be perfectly healthy while a specific **task** for one flow is FAILED, PAUSED, or stuck rebalancing. Health must be assessed **per connector/task**, not per worker. ## Diagnostic order for stale heartbeats 1. **Task state** — query the Connect REST API `GET /connectors/<flow>-MirrorHeartbeatConnector/status`. Look for FAILED (read the trace), PAUSED, or UNASSIGNED tasks. Same for `MirrorSourceConnector`, which is what *replicates* heartbeats to the target. 2. **Source emission** — consume the `heartbeats` topic **on the source**. If new records appear there, the heartbeat connector is fine and the break is in replication; if not, the heartbeat connector/task is the culprit. 3. **Replication leg** — if source heartbeats are fresh but the target's `<source>.heartbeats` is stale, MirrorSourceConnector (or its consumer/producer) is the problem: rebalance loop, target broker unreachable, ACL/authentication failure, or missing topic auto-create permissions on the target. 4. **Target side** — broker reachability, `<source>.heartbeats` topic existence and write ACLs, quota throttling. 5. **Metric absence** — a dead task emits *no* MirrorSourceMetrics, so absence of replication-latency-ms corroborates a stall rather than slowness. ## Distinguishing causes quickly - Source heartbeats fresh + target stale -> replication leg / target connectivity / ACLs. - Source heartbeats also stale -> heartbeat connector/task or source produce issue. - replication-latency-ms present and climbing (not absent) -> not a stall; it's a throughput/backpressure problem (network, target broker capacity, too few tasks). ## Edge cases - **Rebalance storms** in the Connect group can intermittently stall tasks; symptoms flap. - **Clock skew** can make latency *look* huge without a real stall — verify NTP before concluding. - **ACL/auth expiry** (e.g. rotated credentials) breaks the producer to target silently. - An over-aggressive `emit.heartbeats.interval.seconds` change won't cause staleness but changes your detection threshold — re-tune alarms. ## Operational takeaway Treat heartbeat staleness as **page-now / DR-not-ready** and replication-latency-ms as a **trend/SLA** signal. Always inspect per-task status via the Connect REST API rather than trusting worker-level 'up'.

  • Source heartbeats are fresh but the target's <source>.heartbeats is stale. Where is the fault?
    In the replication leg — MirrorSourceConnector's task, or its connection/ACLs/auth to the target broker (or topic auto-create). The heartbeat emitter itself is fine.
  • Why isn't 'the Connect cluster is up' sufficient to declare the flow healthy?
    Workers being alive doesn't mean the specific tasks for that flow are RUNNING; a task can be FAILED, PAUSED, or stuck rebalancing while workers report healthy. Check per-task status.
  • How do you tell a stall from mere slowness using metrics?
    A stall shows heartbeat staleness and often absent replication-latency-ms (dead task emits nothing); slowness shows present-but-rising replication-latency-ms with heartbeats still arriving.

saying these in an interview costs you the question

  • Equating worker-level 'Connect is up' with task/flow health
  • Treating heartbeat staleness as the same thing as high replication-latency-ms
  • Forgetting to check the source heartbeats topic to isolate emit vs replication
  • Assuming absent metrics mean zero/healthy rather than a dead task
  • Overlooking ACL/auth expiry or rebalance storms as stall causes

context