skip to content

Define RPO and RTO for a Kafka DR setup. What practical factors drive each, and what RPO does asynchronous replication imply?

level: middleimportance: must knowfreq 55%

answer

  1. RPO = data loss (time); RTO = downtime (time)
  2. async MM2/Cluster Linking → RPO > 0 ≈ replication lag
  3. RTO driven by detection + client repoint + offset translation
  4. stretched cluster acks=all → RPO≈0 but needs low latency
  5. acks=all is source durability, not cross-region RPO

basics

~20 s

RPO (Recovery Point Objective) = how much data you can lose, measured in time. RTO (Recovery Time Objective) = how long recovery takes. Async cross-cluster replication (MirrorMaker 2) means RPO > 0: anything not yet replicated when the primary dies is lost.

solid answer

~50 s

RPO is the maximum acceptable data loss expressed as a time window — with asynchronous replication it equals roughly the replication lag at the moment of failure, so RPO is always greater than zero. RTO is the maximum acceptable downtime to restore service after a disaster. For Kafka DR, RPO is driven by your replication mechanism and its lag (MirrorMaker 2 / Cluster Linking are async), producer acks, and network throughput between regions. RTO is driven by failover automation: detecting the outage, repointing producers/consumers to the DR cluster, translating consumer offsets, and warming caches. Synchronous replication (a stretched cluster across AZs with acks=all) can give RPO≈0 but needs low latency and quorum across sites, which is impractical across distant regions. Most multi-region Kafka DR is active-passive async, accepting a small RPO in exchange for geographic independence.

go deeper

for a junior

Recall that RPO is how much data you can lose and RTO is how long recovery takes.

for a middle

Connect async replication to a nonzero RPO and list what drives each objective.

for a senior

Quantify RPO from replication lag, design failover to hit RTO, and explain the stretched-cluster RPO≈0 trade-off.

for a principal

Set RPO/RTO SLAs with the business, choose topology (active-passive vs stretched) per latency budget, and own the failover automation.

## The two objectives Disaster recovery is specified by two numbers agreed with the business: - **RPO — Recovery Point Objective**: the maximum amount of data, measured in **time**, you are willing to lose. 'RPO = 5 minutes' means after a disaster you may lose up to the last 5 minutes of writes. It answers *how much data*. - **RTO — Recovery Time Objective**: the maximum **downtime** you tolerate before service is restored on the DR side. 'RTO = 30 minutes' means you must be serving traffic again within 30 minutes. It answers *how long*. ## Why async replication implies RPO > 0 Kafka cross-cluster DR tools — **MirrorMaker 2 (MM2)** and **Cluster Linking** — are **asynchronous**: the source cluster acknowledges a producer before the record reaches the remote cluster. There is always some **replication lag**. If the primary region is destroyed instantly, every record produced but not yet copied is lost. So your effective RPO ≈ the replication lag at failure time. You minimize it with bandwidth, healthy MM2 tasks, and monitoring `replication-latency-ms` / consumer lag on the mirror, but you cannot drive it to zero with async. ## What drives RPO - Replication lag (MM2 task throughput, network RTT/bandwidth between regions). - Producer `acks` (acks=all ensures the record is durable on the *source*, not the remote — it does not reduce cross-region RPO by itself). - Batching/linger and backpressure during traffic spikes. ## What drives RTO - **Detection**: how fast you notice the primary is down (health checks, monitoring). - **Failover mechanics**: repointing clients (DNS, bootstrap.servers, service discovery), and crucially **offset translation** so consumers resume at the right place on the DR cluster (MM2's `RemoteClusterUtils` / `MirrorCheckpointConnector`). - **Automation maturity**: scripted/automated failover vs manual runbooks. - **State warm-up**: rebuilding compacted state, cache priming. ## RPO≈0 alternative A **stretched cluster** (single cluster with replicas/`min.insync.replicas` spread synchronously across AZs, producers using acks=all) achieves RPO≈0 because a write isn't acked until it's on multiple sites. But synchronous quorum needs low inter-site latency, so it suits multi-AZ within a region, not distant multi-region DR. Hence most multi-region designs are **active-passive async** with a small, monitored RPO.

  • Does acks=all give you RPO=0 across regions?
    No. acks=all guarantees durability on the source cluster's in-sync replicas. With async MM2/Cluster Linking the record may not have reached the remote cluster, so cross-region RPO is still the replication lag.
  • How would you architect for RPO≈0?
    A stretched/synchronous cluster spanning AZs with replicas and min.insync.replicas across sites and acks=all, so a write isn't acknowledged until durable on multiple sites. Practical only with low inter-site latency, typically multi-AZ within one region.

saying these in an interview costs you the question

  • Swapping the definitions of RPO and RTO.
  • Claiming async MirrorMaker can deliver RPO=0.
  • Believing acks=all on the producer eliminates cross-region data loss.
  • Ignoring offset translation as part of RTO.

context