Which MM2 JMX metrics would you use to monitor replication lag and latency, and what does each one mean?
answer
- domain kafka.connect.mirror / MirrorSourceMetrics
- replication-latency-ms = source-append to target-append
- record-age-ms = staleness when MM2 reads source
- tags: source/target/topic/partition
- watch -max for SLA; alert on metric absence too
basics
~10 sMM2's MirrorSourceConnector exposes per-topic-partition JMX metrics: replication-latency-ms (time from source append to target append), record-age-ms (age of a record when consumed from source), and byte/record rates. You scrape these to track lag.
solid answer
~40 sMirrorSourceMetrics (MBean domain kafka.connect.mirror) exposes per source-topic-partition gauges. The key ones for lag: `replication-latency-ms` (and its -max/-avg/-min variants) — the time between a record being appended at the source and appended at the target, i.e. true end-to-end replication latency. `record-age-ms` — how old a record was (now minus its source timestamp) at the moment MM2 read it from the source, which surfaces consumer-side lag in the replication pipeline. There are also throughput gauges: `byte-rate`, `record-rate`, and `byte-count`/`record-count`. Tags include source, target, topic, partition. In practice you scrape replication-latency-ms-max for SLA/RPO alarms and combine it with heartbeat-derived end-to-end latency as an independent cross-check. record-age-ms rising signals MM2 is falling behind reading the source; replication-latency-ms rising signals the whole path is slow.
go deeper
Know that MM2 exposes JMX metrics and that replication-latency-ms measures how far behind the target is.
Differentiate replication-latency-ms from record-age-ms and name the JMX domain/tags.
Use the two metrics to localize whether the consume leg or the produce-to-target leg is the bottleneck, and alert on metric absence.
Design a monitoring strategy combining heartbeats + JMX with cardinality control and clock-skew/SLA considerations across the topology.
## Where the metrics live MM2 runs on Kafka Connect, and MirrorSourceConnector publishes metrics under the JMX domain **`kafka.connect.mirror`** as **MirrorSourceMetrics**, tagged by `source`, `target`, `topic`, and `partition`. They are **per source-topic-partition**, so cardinality can be high on big topologies. ## The lag/latency metrics - **`replication-latency-ms`** (with `-avg`, `-max`, `-min`): the elapsed time between when a record was **appended on the source** and when its replicated copy was **appended on the target**. This is the true **end-to-end replication latency** for actual data — the most direct lag signal. Watch `replication-latency-ms-max` for worst-case SLA/RPO. - **`record-age-ms`** (with `-avg`/`-max`/`-min`): when MM2 *reads* a record from the source, this is `consumeTime - recordTimestamp` — how stale the record already was when MM2 picked it up. Rising record-age means MM2's **consumer side is lagging** (it can't keep up reading the source), as opposed to the produce side to the target being slow. - **Throughput**: `record-rate`, `byte-rate`, `record-count`, `byte-count` — volume gauges/counters useful for capacity and for correlating latency spikes with load. ## How they combine for diagnosis - `record-age-ms` high but `replication-latency-ms` not catastrophic -> MM2 is **behind on consuming** the source (under-provisioned tasks, source-side backlog). - `replication-latency-ms` high -> the **full path** is slow (network, target broker throughput, produce backpressure). - Both flat while heartbeat arrivals stall -> likely a **broken/stalled flow** rather than mere slowness — cross-check with the heartbeat topic. ## Heartbeats vs JMX as independent signals Heartbeat-derived latency (`now - heartbeat.timestamp` on the target) is a black-box, data-independent probe that works even on idle topics. The JMX `replication-latency-ms` is a white-box, per-partition measurement on real data. Best practice is to use **both**: heartbeats for liveness/idle-flow coverage and an RPO proxy; JMX for per-topic, per-partition granularity and root-cause. ## Edge cases - Metrics are emitted only while the connector task is running; a dead task means *missing* metrics, not zero — alert on absence too. - Per-partition cardinality can overwhelm a metrics backend on large topologies; aggregate or sample. - Clock skew between source and target brokers distorts both heartbeat latency and replication-latency-ms; keep clusters NTP-synced. - record-age-ms depends on producers setting sensible record timestamps (CreateTime vs LogAppendTime).
- record-age-ms is climbing but replication-latency-ms looks okay — what does that tell you?MM2's consumer side is falling behind reading the source (backlog or under-provisioned tasks), rather than the write-to-target leg being slow.
- Why monitor for the absence of these metrics, not just high values?If a connector task dies, it stops emitting metrics entirely — you see no data rather than a high value, so 'no metric' must itself trigger an alarm.
- How can clock skew corrupt these measurements?replication-latency-ms and heartbeat latency both subtract timestamps taken on different clusters; if source and target clocks drift, the computed latency is wrong — keep both NTP-synced.
saying these in an interview costs you the question
- Treating record-age-ms and replication-latency-ms as the same thing
- Assuming a healthy metric value when the task is dead (you'd see no metric at all)
- Forgetting clock skew distorts latency math across clusters
- Claiming MM2 lag is read from consumer-group lag of the data topics rather than connector JMX