skip to content

Which MM2 JMX metrics would you use to monitor replication lag and latency, and what does each one mean?

level: seniorimportance: must knowfreq 45%

answer

  1. domain kafka.connect.mirror / MirrorSourceMetrics
  2. replication-latency-ms = source-append to target-append
  3. record-age-ms = staleness when MM2 reads source
  4. tags: source/target/topic/partition
  5. watch -max for SLA; alert on metric absence too

basics

~10 s

MM2's MirrorSourceConnector exposes per-topic-partition JMX metrics: replication-latency-ms (time from source append to target append), record-age-ms (age of a record when consumed from source), and byte/record rates. You scrape these to track lag.

solid answer

~40 s

MirrorSourceMetrics (MBean domain kafka.connect.mirror) exposes per source-topic-partition gauges. The key ones for lag: `replication-latency-ms` (and its -max/-avg/-min variants) — the time between a record being appended at the source and appended at the target, i.e. true end-to-end replication latency. `record-age-ms` — how old a record was (now minus its source timestamp) at the moment MM2 read it from the source, which surfaces consumer-side lag in the replication pipeline. There are also throughput gauges: `byte-rate`, `record-rate`, and `byte-count`/`record-count`. Tags include source, target, topic, partition. In practice you scrape replication-latency-ms-max for SLA/RPO alarms and combine it with heartbeat-derived end-to-end latency as an independent cross-check. record-age-ms rising signals MM2 is falling behind reading the source; replication-latency-ms rising signals the whole path is slow.

go deeper

for a junior

Know that MM2 exposes JMX metrics and that replication-latency-ms measures how far behind the target is.

for a middle

Differentiate replication-latency-ms from record-age-ms and name the JMX domain/tags.

for a senior

Use the two metrics to localize whether the consume leg or the produce-to-target leg is the bottleneck, and alert on metric absence.

for a principal

Design a monitoring strategy combining heartbeats + JMX with cardinality control and clock-skew/SLA considerations across the topology.

## Where the metrics live MM2 runs on Kafka Connect, and MirrorSourceConnector publishes metrics under the JMX domain **`kafka.connect.mirror`** as **MirrorSourceMetrics**, tagged by `source`, `target`, `topic`, and `partition`. They are **per source-topic-partition**, so cardinality can be high on big topologies. ## The lag/latency metrics - **`replication-latency-ms`** (with `-avg`, `-max`, `-min`): the elapsed time between when a record was **appended on the source** and when its replicated copy was **appended on the target**. This is the true **end-to-end replication latency** for actual data — the most direct lag signal. Watch `replication-latency-ms-max` for worst-case SLA/RPO. - **`record-age-ms`** (with `-avg`/`-max`/`-min`): when MM2 *reads* a record from the source, this is `consumeTime - recordTimestamp` — how stale the record already was when MM2 picked it up. Rising record-age means MM2's **consumer side is lagging** (it can't keep up reading the source), as opposed to the produce side to the target being slow. - **Throughput**: `record-rate`, `byte-rate`, `record-count`, `byte-count` — volume gauges/counters useful for capacity and for correlating latency spikes with load. ## How they combine for diagnosis - `record-age-ms` high but `replication-latency-ms` not catastrophic -> MM2 is **behind on consuming** the source (under-provisioned tasks, source-side backlog). - `replication-latency-ms` high -> the **full path** is slow (network, target broker throughput, produce backpressure). - Both flat while heartbeat arrivals stall -> likely a **broken/stalled flow** rather than mere slowness — cross-check with the heartbeat topic. ## Heartbeats vs JMX as independent signals Heartbeat-derived latency (`now - heartbeat.timestamp` on the target) is a black-box, data-independent probe that works even on idle topics. The JMX `replication-latency-ms` is a white-box, per-partition measurement on real data. Best practice is to use **both**: heartbeats for liveness/idle-flow coverage and an RPO proxy; JMX for per-topic, per-partition granularity and root-cause. ## Edge cases - Metrics are emitted only while the connector task is running; a dead task means *missing* metrics, not zero — alert on absence too. - Per-partition cardinality can overwhelm a metrics backend on large topologies; aggregate or sample. - Clock skew between source and target brokers distorts both heartbeat latency and replication-latency-ms; keep clusters NTP-synced. - record-age-ms depends on producers setting sensible record timestamps (CreateTime vs LogAppendTime).

  • record-age-ms is climbing but replication-latency-ms looks okay — what does that tell you?
    MM2's consumer side is falling behind reading the source (backlog or under-provisioned tasks), rather than the write-to-target leg being slow.
  • Why monitor for the absence of these metrics, not just high values?
    If a connector task dies, it stops emitting metrics entirely — you see no data rather than a high value, so 'no metric' must itself trigger an alarm.
  • How can clock skew corrupt these measurements?
    replication-latency-ms and heartbeat latency both subtract timestamps taken on different clusters; if source and target clocks drift, the computed latency is wrong — keep both NTP-synced.

saying these in an interview costs you the question

  • Treating record-age-ms and replication-latency-ms as the same thing
  • Assuming a healthy metric value when the task is dead (you'd see no metric at all)
  • Forgetting clock skew distorts latency math across clusters
  • Claiming MM2 lag is read from consumer-group lag of the data topics rather than connector JMX

context